Skip to main content

What to Put in an AI Agent Testing Framework: The Decisions Before You Write a Single Script

Before you pick a tool, four architectural decisions determine whether your AI agent testing framework catches real failures. Here is what they are.

Most teams building an AI agent testing framework reach for tools first. They evaluate libraries, spin up infrastructure, and start writing scripts before they have answered the questions that actually determine whether the framework will catch real failures. The result is a suite that runs green while agents misroute requests, hallucinate tool calls, or silently degrade across multi-turn conversations.

This post argues that an AI agent testing framework is fundamentally an architectural artifact, not a tooling decision. The choices you make before writing a single script, covering which interaction types fall in scope, how failure is defined and observed, and how test results feed back into the product, determine whether your framework functions as a genuine safety net or a false confidence machine.

In the sections that follow, you will work through four foundational decisions, examine how they interact, and see what a framework document looks like before any scripts exist. If your current approach skips straight to implementation, the green suite problem may already be present in your pipeline. Understanding the architecture first is how you close that gap.

The green suite problem

An AI agent completes your checkout flow on every test run. Green across the board. Meanwhile, real users are abandoning at the address autofill step, because the agent under test never navigates that interaction type. Nobody scripted it. The framework never knew it existed.

This is the green suite problem: a framework that reports no failures while real agent failures accumulate in production, not because the scripts are wrong, but because the architectural decisions that determine what gets covered were never made.

Most teams reach for an ai agent testing framework the same way they reach for any tooling: pick an evaluation library, write some scripts, run them in CI. The tool selection happens before anyone has written down what failure means, which interaction types are in scope, or how a detected failure is supposed to reach the team that can fix it. The result is a framework shaped by what the tool makes easy to script, not by what the agent actually does in production. The testing practice gaps documented in open-source AI agent framework research reflect a structural problem, not an isolated one.

The received wisdom is that "ai agent testing" is a tooling problem. It is not. It is an architecture problem. And the architecture has to be decided before a single script is written.

The question this piece works through: what are those decisions, and what breaks if you skip one?

There are four of them. First, interaction scope: the explicit list of agent-to-product interaction types your framework is responsible for covering. Second, failure taxonomy: a written classification of failure modes with definitions precise enough to apply without human interpretation. Third, the observation layer: the combination of signals and diagnostics your framework uses to determine whether a failure has occurred. Fourth, the feedback loop: the mechanism by which a detected failure moves from test result to a product change, with a named owner at each step.

These four decisions are a dependency chain. Scope determines which failure modes are even possible. Taxonomy determines how they are classified. The observation layer determines whether they are detectable. The feedback loop determines whether detection produces any change at all. Skip one, and the decisions downstream of it are built on a gap.

Decision one: which interaction types are in scope

Interaction scope is the explicit list of agent-to-product interaction types your framework is responsible for covering, written down before any script exists. Not implied by the scripts. Written down first.

For web-facing AI agents, that list maps to four categories. Discovery: can the agent locate the thing it needs, navigating menus, search, and linked surfaces? Information retrieval: does it return accurate data from what it finds? Task completion: does it finish a multi-step flow without abandoning mid-sequence? Delegated access: does it behave correctly when acting on behalf of a user, with their credentials and permissions in play? If you want to understand why these four categories matter in practice, what flow AI is and why product teams need to test for it is a useful frame before going further.

Every downstream decision depends on this list. The failure taxonomy and observation layer, covered in the next two sections, are both derivative of it.

The cost of skipping this step is predictable. Teams without an explicit scope document default to testing what is easiest to script, which is almost always happy-path task completion. An agent that books a room, retrieves a document, or completes a purchase gets a test. Discovery and delegated access get nothing, because they require more setup and their failure modes are less obvious. The suite runs green. The real failures go undetected.

Here is a concrete version of that failure. An agent trained to book a meeting room passes every test run: it finds the room, checks availability, and confirms the booking. But when a second agent holds a delegated session on the same calendar with write permissions, the first agent fails to surface the conflict and books anyway. The failure is real. It is not caught because delegated access was never declared in scope, so no script was written for it, and no observation was instrumented around it.

Writing the scope definition takes one sitting. Build a table. Rows are the agent's task types. Columns are the four interaction categories. Each cell is either in-scope or out-of-scope, with a one-line reason. "Delegated access: out-of-scope, single-user deployment only" is a valid entry. "Delegated access: in-scope, agent acts on behalf of authenticated users in calendar and booking flows" is also valid. What is not valid is leaving the cell blank.

Treat this table as a constraint document. If a proposed test script does not map to a row and a column in the table, it should not be written until the table is updated. The table does not grow organically. It grows by decision.

Decision two: what counts as a failure

With your scope table written, you need a second document to sit beside it: a failure taxonomy. A failure taxonomy is a written classification of the failure modes your framework is able to detect, where every definition is stated in observable terms that do not require human interpretation to apply. If two engineers reading the same definition would produce different pass/fail logic, the definition is not finished.

The three broad categories for AI agents in product contexts are:

  • Behavioral failures: the agent does the wrong thing. It navigates to the wrong page, returns an incorrect value, or takes an action outside its declared task.

  • Performance failures: the agent does the right thing too slowly or inconsistently. The correct result arrives after a threshold that makes it unusable in context.

  • Reliability failures: the agent stops mid-task, or produces materially different output across identical inputs.

Shared definitions must precede scripting. Without them, two engineers writing tests for the same agent will apply different pass/fail logic to identical behavior, results cannot be aggregated, and trend data across runs becomes noise.

The non-determinism problem

Binary pass/fail logic, inherited from unit testing, is the wrong instrument for AI agents. An agent can return different outputs on identical inputs without any underlying fault: model temperature, retrieved context, and session state all introduce variance. The taxonomy must specify acceptable variance ranges per failure mode and define the threshold at which deviation becomes a failure signal. A behavioral failure in an information retrieval task, for example, might be defined as: the returned value deviates from the ground-truth value by more than a named tolerance across three consecutive runs on the same input.

This is why what Meta AI actually does to your website matters beyond advertising: external AI systems acting on your product create variance your taxonomy needs to account for, because their behavior affects what your own agent sees and retrieves.

Agent failure versus environment failure

A test that fails because a third-party API timed out is not an agent failure. Conflating the two inflates false positive rates until the team stops trusting the suite. Every failure mode in your taxonomy needs a detection method that distinguishes between the agent's behavior and the environment's behavior. The arXiv taxonomy of faults in agentic AI systems separates LLM integration faults from lifecycle and state failures for exactly this reason: the root cause determines the fix, and the fix determines which team owns the response.

What to produce

The output of this decision is a document with named failure modes. Each entry contains: a name, a definition in observable terms, the interaction type it maps to from your scope table, and the detection method. A compact taxonomy, enough modes to cover each interaction type without redundant overlap, outperforms an exhaustive one where definitions require judgment calls to apply.

Decision three: how you will observe the agent

A failure taxonomy tells you what you are looking for. The observation layer is how you actually see it.

The observation layer is the combination of signals, logs, and diagnostics your framework uses to determine whether a failure mode from your taxonomy occurred during a test run. Without it, the taxonomy is a list of things you have decided to care about but cannot detect.

Three approaches, three trade-offs

Black-box observation treats the agent as opaque: you measure inputs and outputs at the product boundary and nothing in between. It is fast to implement and works for output-level failure modes like wrong return value or task abandonment. It cannot tell you why the failure happened.

White-box observation instruments the agent's internal reasoning: every tool call, retrieval step, and decision node is logged. You can trace a failure back to the exact step where the agent's reasoning diverged. The cost is data volume and the requirement that you have access to the agent's internals.

Hybrid observation applies black-box measurement at the product boundary and white-box tracing for the agent's decision path. It is the right default for most AI agent testing contexts because it separates product-level failures from agent-level failures without requiring you to instrument everything.

The reproducibility problem

An observation layer is only as reliable as the environment it runs in. Traditional static test environments do not simulate autoscaling, live session state, or third-party dependency behaviour under load. A failure that only appears when two agent sessions hold concurrent access to the same resource will not appear in a static environment, because the concurrency condition is never reproduced. You build a framework that cannot see the failures that matter most in production.

What to instrument, at minimum

Four things: the full request/response trace at the product boundary; the agent's action sequence, meaning each tool call or navigation step in order; any confidence scores or retrieval results the agent used to make decisions; and the wall-clock time for each step. These four signals let you confirm a failure occurred, locate where in the action sequence it happened, and check whether a slow step or a low-confidence retrieval preceded it.

The gap in current practice

Most teams capture output and stop there. They can tell that a failure occurred. They cannot tell which reasoning step produced it, so the fix is a guess. McKinsey's research on AI trust named lack of trace-level visibility as one of the top reasons agent rollouts stall, and that gap shows up here: detection without a reasoning trace produces a dashboard of failures that no one can act on confidently.

This is the distinction that matters when you are thinking about what it means for AI agents to actually use your product: detecting that an agent failed a task is different from knowing where in the agent's navigation path the failure originated.

Stunt Double sends AI agents through a product to simulate how real users and agents experience it, which produces a hybrid observation layer by default: product-boundary outputs alongside the agent's navigation path across discovery, retrieval, and task completion.

A note on data volume

White-box instrumentation of every reasoning step produces large trace volumes quickly. The observation layer architecture must decide before any scripts run what is stored in full, what is sampled, and what is discarded. Without that decision, storage costs scale with test coverage in a way that makes expanding the suite expensive, which is the wrong incentive.

Decision four: how test results reach the product

Detection without routing is a filing system, not a feedback loop.

The feedback loop is the mechanism by which a detected failure moves from test result to a product change: a named owner at each step, and a defined latency from signal to resolution. Without those two properties written into the framework, the path defaults to manual review. Manual review means someone checks the dashboard when they remember to. Failures accumulate. The suite keeps running. Nothing changes.

This is an architectural decision, not a process decision. A process decision is a team norm: "we review test results on Fridays." An architectural decision is a structural constraint: the framework cannot emit a failure signal without also emitting a routing target. The difference is whether the loop closes automatically or relies on a person to notice.

Evaluation frameworks typically surface that a failure occurred. The gap, which has no standardized solution in current practice, is the translation layer between a failure signal and a product change. That translation layer is what determines whether the framework changes anything at all.

The four components that close the loop:

  • Routing rule: each failure mode maps to a named team. A behavioral failure routes to the team that owns the agent's decision logic. A performance failure routes to the team that owns the product's integration layer. The rule is written down before the first test runs, not inferred from whoever reads the report.

  • Severity threshold: defines what triggers an immediate escalation versus what enters the backlog. A threshold specific enough to act on names a failure rate and a named recipient. "Serious failures get escalated" is not specific enough.

  • Retest trigger: when a fix is shipped, the test that detected the original failure reruns against the changed product. Not a new test. The same test. This confirms the fix resolves the specific condition that triggered the signal.

  • Trend signal: the framework surfaces patterns across runs, not just individual failures. A single failure on one run is noise. The same failure recurring across runs over weeks is a product problem.

The concrete cost of omitting this: a framework detects that an agent abandons information retrieval tasks at a specific page on a significant share of runs. There is no routing rule. The signal sits in a test report. The page is not changed. The framework detects the same failure on the next run, and the one after that. The detection is working. Nothing else is.

The metric that surfaces a broken loop: mean time between a failure's first detection and the product change that resolves it. If there is no defined maximum latency between a failure's first detection and the product change that resolves it, the loop is broken by design. This is where experience decay after launch goes unmeasured: not because no one is watching, but because the path from signal to fix was never designed.

How the four decisions interact

The dependency chain introduced earlier has a predictable failure cascade worth mapping in full.

A team that builds a solid taxonomy and instruments a capable observation layer, but never writes an explicit scope definition, will have high detection fidelity on every interaction type they happened to script and zero coverage on every interaction type they never thought to include. The suite runs green. Delegated access failures, discovery failures, the gaps accumulate silently. The problem is not the scripts. The scripts are fine. The problem is that scope was never declared, so the scripts define it by default.

Tooling compounds this. Any evaluation library or testing tool selected before these four decisions are made will shape scope implicitly: you will test what the tool makes easy to test. Output comparison tools bias toward task completion. Latency monitors bias toward performance. The interaction types that do not fit cleanly into the tool's model get deprioritized, then dropped, then forgotten. This is how an AI agent's actual behavior through your product diverges from what the framework covers: not through negligence, but through a series of small tool-shaped choices that were never recognized as scope decisions.

Environment complexity cuts across all four decisions. The testing environment must simulate the conditions under which real failures occur, including autoscaling behavior, session concurrency, and third-party dependency states. Manual environment setup and teardown consumes over 40% of testing cycles in practice. That cost is not an operational inconvenience: it is what happens when the environment is not designed into the architecture from the start. Each of the four decisions requires a stable, realistic environment to produce reliable results. Without it, the observation layer captures data from conditions that do not match production, the taxonomy fires on environment noise rather than agent behavior, and the feedback loop routes false signals to engineering teams who learn to ignore them.

The final cross-cutting concern is drift. Agent behavior changes when the underlying model updates, even when no one on the team made a deliberate change. A framework built for last quarter's agent scope will miss this quarter's failure modes. Each of the four decisions should be revisited when agent capabilities change, not only when a new failure is detected. Treat it as a reroll: same decisions, updated inputs, new constraints documented.

What the framework document looks like before any scripts exist

The four decisions produce a single document. Four sections, nothing more: an interaction scope table, a failure taxonomy with detection methods, an observation layer specification, and a feedback loop diagram with named owners and latency targets. The scripts implement this document. The document is the framework.

The scope table

The scope table is a grid. Rows are agent task types: book a room, retrieve a policy, complete a purchase. Columns are the four interaction categories: discovery, retrieval, task completion, delegated access. Each cell holds one of two values, in-scope or out-of-scope, followed by a single line of reasoning. "Delegated access / book a room: out of scope, no third-party calendar integration in current release." That line is doing real work. It means no engineer writes a delegated-access script for room booking until someone updates that cell and the reason behind it.

A scope table with fifteen rows and four columns is achievable in a single working session. If it takes much longer, the team does not yet understand its own agent's task surface well enough to write scripts.

The failure taxonomy

Each row in the taxonomy names a failure mode, defines it in observable terms (what the agent did or did not do, not what it "seemed to intend"), assigns it to an interaction type from the scope table, and specifies a detection method. Detection methods should be drawn from a constrained set, such as output comparison, timing threshold, or action sequence check, rather than left free-form. Understanding what an AI generator does when it visits your product informs which detection method applies: an agent navigating your product produces an action sequence, and sequence-check failures are distinct from output failures.

A compact taxonomy, enough modes to cover each interaction type without redundant overlap, outperforms an exhaustive one where definitions require judgment calls to apply.

The timeline and the gate

Both sessions together are scoped to complete before any scripting begins, treat them as a prerequisite, not a parallel track. Session one: the scope table, agreed and signed off. Session two: the taxonomy, observation layer specification, and feedback loop diagram. The observation layer specification is a list, not prose: what is instrumented, the signal type for each item, and the retention policy. The feedback loop diagram names each hand-off point, the owner at that point, and the maximum acceptable latency before escalation.

Once the document exists, it creates a gate. A proposed test script that cannot be traced to a named interaction type in the scope table and a named failure mode in the taxonomy does not get written. The document gets updated first, or the script does not exist. This is not bureaucratic friction. It is the mechanism that keeps the framework from drifting back toward testing only what is easy to script.

Before the first script

That document, once written, is the framework. The scripts are its implementation.

The practical sequence is flat: write the scope table before opening any test runner; define each failure mode in observable terms before selecting any tool, because the tool you choose will otherwise define your failure modes for you; instrument the observation layer to capture the agent's reasoning path, not just its final output; and assign named owners to the feedback loop before the first test runs, not after the first failure report lands with no clear recipient.

Each decision has one quality check. After writing it, ask whether someone who was not in the room when it was written could apply it independently. If the scope table requires oral context to interpret, it is not specific enough. If a failure mode says "clearly wrong output," it is not observable. If the feedback loop says "escalate as needed," it has no owner. The check is not strict for its own sake: definitions that fail it will produce inconsistent pass/fail logic across engineers, and aggregated results from inconsistent logic are noise.

The check applies to ai agent testing of any kind. Evaluation frameworks that skip it produce the same outcome: detection theater, green suites, real failures accumulating in production while the dashboard shows passing runs.

Most frameworks do not fail because the scripts are wrong. They fail because the decisions that should have shaped the scripts were never written down.

Conclusion

A reliable AI agent testing framework is not built with scripts. It is built with decisions made before a single script exists.

Write each decision down. Apply the independence check. If someone outside the room cannot apply it without explanation, rewrite it.

Start with the scope table today. Everything useful in your testing infrastructure follows from it.