teamcity-ai-agent-testing-demo — Gemini AI Agent
    Neura MarketNeura Market/Gemini
    ChatGPTChatGPTClaudeClaudeGeminiGeminiCursorCursorGrokGrokPerplexityPerplexityDeepSeekDeepSeek
    CoPilotCoPilotStable DiffusionStable DiffusionMidjourneyMidjourney
    View All Directories
    OverviewRulesPromptsMCPsAgentsGamesBlogVideosGuidesCoursesCommunityGemsExtensionsTrending
    GeminiAgentsteamcity-ai-agent-testing-demo
    Back to Agents
    teamcity-ai-agent-testing-demo

    teamcity-ai-agent-testing-demo

    JetBrains July 1, 2025
    5 copies 0 downloads

    End-to-end TeamCity framework to run AI agents on SWE-Bench Lite. Spin up isolated Docker images per task, extract patches, score with the official harness, and aggregate success rates. As an example, we'll look at Junie and Google Gemini CLI

    Image

    <div align="center"> <p> <a href="#overview">Overview</a> • <a href="#project-structure">Project structure</a> • <a href="#setup-instructions">Setup instructions</a> • <a href="#monitoring-and-results">Monitoring and results</a> • <a href="https://www.jetbrains.com/teamcity/use-cases/ai/" target="_blank">TeamCity for AI agent evaluation ↗️</a> </p> </div>

    SWE-Bench AI Agent Testing with TeamCity

    TeamCity gives you a reproducible CI/CD backbone for AI agent evaluation: orchestrate parallel, isolated (Docker) runs across benchmark tasks (e.g., SWE-Bench), validate patches automatically, and capture metrics, logs, and artifacts to track performance and costs at scale via Kotlin-DSL pipelines.

    This TeamCity configuration provides a complete framework for testing AI agents against the SWE-Bench Lite dataset, which contains 300+ software engineering tasks from popular Python repositories.

    Overview

    The system evaluates AI agents by:

    1. Preparing isolated Docker environments for each SWE-Bench task
    2. Running AI agents against specific coding problems
    3. Evaluating solutions using the official SWE-Bench evaluation harness
    4. Collecting performance metrics and success rates

    Architecture

    Projects Structure

    • JetBrains Junie AI Agent (JetBrain_Junie_AI_Agent.kt)

      • Downloads Junie CLI from GitHub releases and IntelliJ IDEA
      • Creates task subsets for progressive testing
      • Individual task execution builds for all 300+ SWE-Bench tasks
    • Google Gemini CLI AI Agent (Google_Gemini_CLI_AI_Agent.kt)

      • Builds from the official Google Gemini CLI repository (https://github.com/google-gemini/gemini-cli.git)
      • Uses Node.js execution environment with npm build process
      • Creates task subsets for progressive testing
    • SWE-Bench Lite (SWE_Bench_Lite.kt): Core dataset and e

    Tags

    agent-evaluationagentic-aiaievalevaluationevaluation-frameworkevaluation-tools

    Comments

    More Agents

    View all
    Agentsmithagentic-ai

    Agentsmith

    Universal, model-agnostic operating harness for AI agents (Claude, Codex, Gemini, …) — a lean core + work-type profiles assembled by one setup script.

    P
    PromptPartner
    308
    Awesome Gamedev Agent Skillsagent-skills

    Awesome Gamedev Agent Skills

    Game-development Agent Skills for AI coding agents: install once and a master router loads the right skill for your engine and task. 66 original, version-pinned skills (plus a master router) in the portable SKILL.md format that runs across Claude Code, Cursor, Codex, Copilot, Gemini CLI and more, for Godot, Unity, Unreal, web and beyond.

    G
    gamedev-skills
    303
    Agentpetai-agents

    Agentpet

    A desktop pet for macOS & Windows that monitors your AI coding agents (Claude Code, Codex, Cursor, Gemini...) in real time, and grows as you code, feed it tokens, level it up, climb the leaderboard.

    N
    ntd4996
    279
    UltraGameStudioai-agent

    UltraGameStudio

    UltraGameStudio - AI coding agent for game development: engine workflows, gameplay code, and asset generation.

    W
    wellingfeng
    260
    Zeroai

    Zero

    The coding agent that answers to you, your model, your machine, your rules.

    G
    Gitlawb
    1,099
    Lucarneagent-bridge

    Lucarne

    Stop babysitting local AI agents. Just notifications, approve, and resume your Codex,Pi,Grok, or Claude code sessions anywhere. 0-Intrusion mobile control bridge via Telegram/微信/飞书. No hooks, no skills, no MCP.

    T
    tuchg
    314

    Stay up to date

    Get the latest Gemini prompts, rules, and resources delivered to your inbox weekly.

    Neura Market LogoNeura Market

    Discover the best AI prompts, plugins, and resources for Gemini and more.

    Content Types

    • Rules
    • Prompts
    • MCPs
    • Agents
    • Guides

    Platforms

    • ChatGPT Directory
    • Claude Directory
    • Gemini Directory
    • Cursor Directory
    • Grok Directory
    • Perplexity Directory
    • DeepSeek Directory
    • CoPilot Directory
    • Stable Diffusion Directory
    • Midjourney Directory
    • All Directories

    Resources

    • Blog
    • Documentation
    • Help Center
    • Marketplace

    Legal

    • Privacy Policy
    • Terms of Service

    © 2026 Neura Market. All rights reserved.

    |

    Not affiliated with any AI platform vendors.

    Neura Market

    Custom AI Systems & Services

    Our team of experienced AI builders will help build custom AI systems, workflows, and solutions for your business.

    Request custom work

    Ready-made automations for this

    Workflows from the Neura Market marketplace related to this Gemini resource

    • Automate Facebook Ad Copy Creation Using Google Sheets and ChatGPT with FAB Frameworkmake · $3.99 · Related topic
    • Generate Facebook Ad Copy Using Google Sheets and ChatGPT with Hero's Journey Frameworkmake · $3.99 · Related topic
    • Load and Aggregate Files from a Google Drive Folder into a Key-Value Dictionaryn8n · $4.99 · Related topic
    • Generate High-Conversion Sales Copy Using Hormozi Framework and Google Docsn8n · $24.99 · Related topic
    Browse all workflows