p.enthalabs

GitHub - gojiplus/understudy: Scenario Testing for AI Agents

![Image 1: PyPI version](https://badge.fury.io/py/understudy)![Image 2: PyPI Downloads](https://pepy.tech/project/understudy)![Image 3: Python 3.12+](https://www.python.org/downloads/)![Image 4: Documentation](https://gojiplus.github.io/understudy/)![Image 5: License: MIT](https://opensource.org/licenses/MIT)

Understudy is a scenario-driven testing framework for AI agents that simulates realistic multi-turn users, runs those scenes against an agent through a simple app adapter, records a structured execution trace of messages, tool calls, and handoffs, and then evaluates behavior with deterministic checks, optional LLM judges, and run reports.

How It Works

[](https://github.com/gojiplus/understudy#how-it-works) Testing with understudy is **4 steps**:

1. **Wrap your agent** — Adapt your agent (ADK, LangGraph, HTTP) to understudy's interface 2. **Mock your tools** — Register handlers that return test data instead of calling real services 3. **Write scenes** — YAML files defining what the simulated user wants and what you expect 4. **Run and assert** — Execute simulations, check traces, generate reports

The key insight: **assert against the trace, not the prose**. Don't check what the agent said—check what it did (tool calls).

Two Evaluation Paradigms

[](https://github.com/gojiplus/understudy#two-evaluation-paradigms)

Conversational Agent Evaluation

[](https://github.com/gojiplus/understudy#conversational-agent-evaluation) Simulate multi-turn conversations with personas to test dialogue agents.

- Use case: Customer service bots, assistants, chatbots

- Assert on tool calls: `trace.called("tool_name")`

Agentic Flow Evaluation

[](https://github.com/gojiplus/understudy#agentic-flow-evaluation) Evaluate autonomous agents executing multi-step tasks.

- Use case: Code agents, research agents, task automation

- Assert on actions: `trace.performed("action")`

See examples/README.md for complete examples of both paradigms.

**See real examples:**

- Example scene — YAML defining a test scenario

- ADK test file — pytest assertions against traces

- LangGraph test file — same tests, different framework

- Agentic test file — agentic flow evaluation

- Agentic scene — task-based scenario

- Example report — HTML report with metrics and transcripts

Installation

[](https://github.com/gojiplus/understudy#installation)

pip install understudy[all]

Quick Start

[](https://github.com/gojiplus/understudy#quick-start)

1. Wrap your agent

[](https://github.com/gojiplus/understudy#1-wrap-your-agent)

from understudy.adk import ADKApp from my_agent import agent

app = ADKApp(agent=agent)

2. Mock your tools

[](https://github.com/gojiplus/understudy#2-mock-your-tools) Your agent has tools that call external services. Mock them for testing:

from understudy.mocks import MockToolkit

mocks = MockToolkit()

@mocks.handle("lookup_order") def lookup_order(order_id: str) -> dict: return {"order_id": order_id, "items": [...], "status": "delivered"}

@mocks.handle("create_return") def create_return(order_id: str, item_sku: str, reason: str) -> dict: return {"return_id": "RET-001", "status": "created"}

3. Write a scene

[](https://github.com/gojiplus/understudy#3-write-a-scene) Create `scenes/return_backpack.yaml`:

id: return_eligible_backpack description: Customer wants to return a backpack

starting_prompt: "I'd like to return an item please." conversation_plan: | Goal: Return the hiking backpack from order ORD-10031.

- Provide order ID when asked

- Return reason: too small

persona: cooperative max_turns: 15

expectations: required_tools:

- lookup_order

- create_return

forbidden_tools:

- issue_refund

4. Run simulation

[](https://github.com/gojiplus/understudy#4-run-simulation)

from understudy import Scene, run

scene = Scene.from_file("scenes/return_backpack.yaml") trace = run(app, scene, mocks=mocks)

assert trace.called("lookup_order") assert trace.called("create_return") assert not trace.called("issue_refund")

Or with pytest (define `app` and `mocks` fixtures in conftest.py):

pytest test_returns.py -v

Suites and Batch Runs

[](https://github.com/gojiplus/understudy#suites-and-batch-runs) Run multiple scenes with multiple simulations per scene:

from understudy import Suite, RunStorage

suite = Suite.from_directory("scenes/") storage = RunStorage()

Run each scene 3 times and tag for comparison

results = suite.run( app, mocks=mocks, storage=storage, n_sims=3, tags={"version": "v1"}, ) print(f"{results.pass_count}/{len(results.results)} passed")

Simulation and Evaluation

[](https://github.com/gojiplus/understudy#simulation-and-evaluation) Understudy separates simulation (generating traces) from evaluation (checking traces). Use together or separately:

Combined (most common)

[](https://github.com/gojiplus/understudy#combined-most-common)

understudy run \ --app mymodule:agent_app \ --scene ./scenes/ \ --n-sims 3 \ --junit results.xml

Separate workflows

[](https://github.com/gojiplus/understudy#separate-workflows) Generate traces only:

understudy simulate \ --app mymodule:agent_app \ --scenes ./scenes/ \ --output ./traces/ \ --n-sims 3

Evaluate existing traces:

understudy evaluate \ --traces ./traces/ \ --output ./results/ \ --junit results.xml

Python API:

from understudy import simulate_batch, evaluate_batch

Generate traces

traces = simulate_batch( app=agent_app, scenes="./scenes/", n_sims=3, output="./traces/", )

Evaluate later

results = evaluate_batch( traces="./traces/", output="./results/", )

CLI Commands

[](https://github.com/gojiplus/understudy#cli-commands)

Run simulations

understudy run --app mymodule:app --scene ./scenes/ understudy simulate --app mymodule:app --scenes ./scenes/ understudy evaluate --traces ./traces/

View results

understudy list understudy show <run_id> understudy summary

Compare runs by tag

understudy compare --tag version --before v1 --after v2

Generate reports

understudy report -o report.html understudy compare --tag version --before v1 --after v2 --html comparison.html

Interactive browser

understudy serve --port 8080

HTTP simulator server (for browser/UI testing)

understudy serve-api --port 8000

Cleanup

understudy delete <run_id> understudy clear

LLM Judges

[](https://github.com/gojiplus/understudy#llm-judges) For qualities that can't be checked deterministically:

from understudy.judges import Judge

empathy_judge = Judge( rubric="The agent acknowledged frustration and was empathetic while enforcing policy.", samples=5, )

result = empathy_judge.evaluate(trace) assert result.score == 1

Built-in rubrics:

from understudy.judges import ( TOOL_USAGE_CORRECTNESS, POLICY_COMPLIANCE, TONE_EMPATHY, ADVERSARIAL_ROBUSTNESS, TASK_COMPLETION, )

Report Contents

[](https://github.com/gojiplus/understudy#report-contents) The `understudy summary` command shows:

- **Pass rate** — percentage of scenes that passed all expectations

- **Avg turns** — average conversation length

- **Tool usage** — distribution of tool calls across runs

- **Agents** — which agents were invoked

The HTML report (`understudy report`) includes:

- All metrics above

- Full conversation transcripts

- Tool call details with arguments

- Expectation check results

- Judge evaluation results (when used)

Documentation

[](https://github.com/gojiplus/understudy#documentation) See the full documentation for:

- Installation guide

- Writing scenes

- ADK integration

- LangGraph integration

- HTTP client for deployed agents

- API reference

License

[](https://github.com/gojiplus/understudy#license) MIT