Agent Identity Is Coming: How to Test Your Auth Flows Before the Standards Land
AI agent authentication testing guide: Learn why agent auth breaks differently, map emerging standards, and build a compatibility checklist before
Your authentication stack was built for humans. Humans type passwords, click "allow," and give up after two failed attempts. AI agents do none of those things, and that mismatch is already causing real problems in production systems handling agentic workflows.
IETF and OpenID Foundation working groups are circulating drafts on verifiable agent credentials right now. No ratified protocol is expected before late 2027, but the trajectory is clear: agent identity will become a baseline expectation, and teams who wait for the ink to dry before testing will face forced migrations under pressure rather than deliberate remediation.
This guide is about getting ahead of that curve through structured ai agent testing against the emerging draft standards. You will learn why agent auth breaks in ways that human auth testing never exposed, where the specifications actually stand today (including which drafts have already expired), and how to run a five-layer compatibility check against your existing flows. By the end, you will have a documented gap register, not a finished implementation, but a clear picture of what needs to change before compliance becomes compulsory.
Why agent auth breaks differently from human auth
An AI agent hits a 401, retries with the same scoped token three times on exponential backoff, and gets the client ID permanently blocked. A human would have stopped at the first prompt, re-authenticated, and continued. The failure mode is different in kind, not degree.
Current Internet protocols conflate authentication and authorisation because they were designed for human-operated clients: a person logs in, consents to a scope, and the session is synchronous. Agents present credentials programmatically, scope permissions to a specific task, and operate across asynchronous sessions where no user is present to respond to a prompt. That mismatch is structural.
OAuth 2.1 handles single trust domains and synchronous operations competently. The OpenID Foundation's own analysis notes it "may fall short in scenarios that are cross-domain, highly autonomous, or asynchronous" (OpenID Foundation, 2025). Most agent deployments do not stay inside a single trust domain for long.
Most small-business deployments compound this by using shared credentials rather than per-agent scoped identities. A shared credential cannot be revoked cleanly: revoking it breaks every service that holds it. That blast-radius problem never surfaces in a human login model because humans do not share credentials at scale.
The 51% of organisations that lack clear ownership of AI identities (WEF, 2025) describes a governance posture built for humans.
Where the standards actually are right now
The governance problem exists partly because the standards meant to fix it do not yet exist. No IETF protocol specification for AI agent identity is expected before December 2027: as of September 2026, every AI agent identity draft is an individual submission with no formal standing, and two of the four most-cited have already expired.
Three active drafts mark the current frontier. draft-klrc-aiagent-auth proposes reusing WIMSE architecture and OAuth 2.0. draft-ni-wimse-ai-agent-identity extends OAuth 2.0 delegation with SPIFFE-style workload identifiers. draft-sharif-agent-identity-framework introduces a five-layer model covering identity, authorisation, attestation, evidence, and trust: the checklist used in the steps below. All three converge on short-lived, audience-restricted tokens that cryptographically bind an agent's identity to the principal that authorised it.
One baseline is already ratified. The Model Context Protocol made OAuth 2.1 and RFC 9728 (OAuth 2.0 Protected Resource Metadata) mandatory as of its June 2025 specification revision, then elevated OAuth Client ID Metadata Documents to recommended status in November 2025. Understanding what flow AI means for product teams clarifies why this matters: agents navigating your product today are already presenting credentials against this spec. AWS, Azure, and GCP all implement SPIFFE/WIMSE natively, so you can test standards-aligned patterns against real infrastructure now. The question of whether your product is legible to AI agents at all is prior to any auth question: a flow an agent cannot parse will never reach your token endpoint.
What you need before you start
With the standards landscape mapped, the next question is practical: what needs to be in place before the first test runs.
An OAuth 2.0 test suite covering at least the authorisation code flow and client credentials grant. If this does not exist, the agent-identity gap is the second problem, not the first. Build the baseline before testing anything on top of it.
A way to replay agent requests through your auth stack. A proxy works. A dedicated test client works. An AI agent testing platform such as Stunt Double works: it can run a scripted actor presenting credentials programmatically rather than through a browser session, which is the credential pattern that matters here.
Access to at least one environment with native SPIFFE/WIMSE support, or a local SPIRE server running in a test namespace. Major cloud providers already implement this natively. This is the workload identity layer the drafts assume; without it, attestation tests have nowhere to land.
RFC 9728 and draft-sharif-agent-identity-framework-01 bookmarked. The five-layer model from the latter serves as the compatibility checklist in the steps below.
An inventory of every non-human identity in your system, with the last rotation date for each credential. Check the rotation date on each result: credentials older than 12 months with no rotation event are candidates for immediate scoping work.
Step 1: Baseline your flow against the MCP OAuth 2.1 mandate
RFC 9728 is the one fixed point in this landscape: ratified, deployed, and already the de facto standard for connecting language models to external tools via MCP. Start here.
Check your authorisation server against three OAuth 2.1 requirements:
PKCE is required on all authorisation code flows
Implicit grant must not be used (PKCE-based authorization code or client credentials only) (per OAuth 2.1 consolidation)
Refresh tokens must be sender-constrained
Each of these is a testable assertion, not a configuration note. Run them against your server now and record the result.
Check your Protected Resource Metadata endpoint. Issue a GET to /.well-known/oauth-protected-resource and confirm the response includes the scope definitions your agents will request. A missing endpoint returns no error to the caller: the agent simply cannot discover what it is allowed to request, and the flow fails silently. A stale endpoint is the same failure with older data.
Validate your OAuth Client ID Metadata Documents. These were elevated to recommended status in the November 2025 revision. Add an automated assertion against this document to your deployment pipeline: it should be machine-readable on every release, not only when someone remembers to check.
Record every gap as a named line in your test log with the relevant RFC 9728 section reference attached. This log is the migration backlog the later steps will build on.
Step 2: Write a test actor that behaves like an agent, not a human
A test actor is a scripted client that presents credentials programmatically, requests the narrowest scope the task requires, and retries on 401 before surfacing a failure. That pattern is what your auth stack will encounter from a production agent, and what your tests must replicate.
Use client credentials grant, not authorisation code. Autonomous agents have no user present to approve an OAuth redirect. Testing with authorisation code flow produces a false pass: the test harness supplies the human approval step the real agent cannot. Configure the actor as a confidential client using client credentials grant only.
Set the actor to retry on 401 with exponential backoff, three attempts maximum, then log a hard failure. Most rate-limiting and lockout policies were tuned for human retry cadences. Programmatic clients hit the threshold faster and trigger lockouts that human-oriented testing never surfaces.
Run the actor against each protected resource in scope and record four things: the token issued, the scopes granted versus the scopes requested, the response time, and whether the token is sender-constrained as OAuth 2.1 requires.
Then compare scopes granted against the minimum needed for the task. Any overage is a blast-radius risk, the same problem AI generators visiting your product expose when they arrive with broader permissions than the task warrants. Stunt Double's delegated-access testing scenes automate this comparison across multiple resource endpoints. Given how much unverified AI-generated code reaches production, running this check on every deployment is not overcaution.
Step 3: Identify where shared credentials are still in use
The blast-radius comparison in the previous step assumes each credential maps to exactly one service. Most do not.
A shared credential is any client ID used by more than one agent or service without a scoped identity. It is the most common agent identity anti-pattern in current deployments and the hardest to revoke cleanly, because revoking it breaks everything that depends on it simultaneously.
Query your identity provider for all non-human client IDs and map each to every service or agent that references it. Any client ID appearing in more than one context is a shared credential by definition. Most identity providers can export this as a service account or application list; the mapping work is manual but bounded.
Check the rotation date on each result. Anything older than 12 months with no rotation event is a candidate for immediate scoping work, not a future backlog item.
For each shared credential found, write a revocation test. Assert that the protected resource returns 401 after the credential is revoked. If more than one service fails, the blast radius is confirmed and logged. The test does not fix the problem; it names the scope of it.
If you are using enterprise SSO with SCIM provisioning for agent lifecycle management, this inventory is your input. Every shared credential is a provisioning gap: an agent the provisioning workflow never created individually and therefore cannot remove cleanly.
Step 4: Run the five-layer compatibility check
Using the five-layer model from draft-sharif-agent-identity-framework-01, each layer is a question. Answer it, then log whether your current flow can or cannot.
Identity layer: can your system issue a stable, unique identifier to an agent, distinct from the human or service account that spawned it? Test against the SPIFFE-style workload identifier pattern in draft-ni-wimse-ai-agent-identity. If your agent inherits the parent service account's identity, the answer is no. Log it.
Authorisation layer: does your authorisation server support downscoped delegation, where an agent acts on behalf of a user but holds narrower permissions? Request a delegation token via draft-ni-wimse-ai-agent-identity's OAuth 2.0 extension. The test is binary: it either produces a downscoped token or it does not. Log the gap if not.
Attestation layer: can your system verify that the agent presenting the credential is the agent it claims to be, rather than a copied or replayed credential? This is where SPIFFE/WIMSE workload attestation applies. Most current flows have no answer here. Log it and move on.
Evidence layer: do your logs record scope-at-use-time, not just token issuance? If an agent operates the way Meta's crawlers do across web properties, acting autonomously at scale with no human in the loop, token issuance alone tells you nothing when something goes wrong. Confirm your logs capture: actor identity, delegating principal, scope asserted, and resource accessed.
Trust layer: if an agent operates cross-domain or spawns a sub-agent, can you verify the delegation chain? Recursive delegation is out of scope in current drafts. A gap here is a known unknown, not a failure. Log it as such.
Step 5: Log every gap as a migration item, not a blocker
The five-layer check produces findings. This step turns them into a log you can act on.
The goal is a named, prioritised list of gaps recorded while there is still time to address them without a forced migration.
Structure each entry with five fields:
Layer: which of the five layers the gap belongs to
Current behaviour: what your flow does today
Expected behaviour: what the relevant draft or standard specifies
Rework effort: an engineering estimate, even a rough one
Priority: based on blast radius if the gap is exploited
Set priority by standing, not severity alone. RFC 9728 gaps are P1 -- the ratified MCP baseline -- and should clear the next sprint. Draft-level gaps are P2 or P3, weighted by proximity to working group adoption. WIMSE is already a chartered working group; drafts in its orbit sit closer to P2 than an individual submission expiring in six months.
Review the log quarterly against IETF Datatracker. When a draft moves from individual submission to working group adoption, its associated gap entries move up in priority. That transition is the signal, not the ratification date.
Share the log with your security and platform teams now. Name the problem: this log gives your security team vocabulary for a governance gap most organisations have not yet formalised.
The failure modes to document, not solve today
Four scenarios belong on your log as named gaps rather than active work items.
Cross-domain agent authorisation. When an agent operates across two organisations with separate identity providers, OAuth 2.1 has no standardised mechanism to handle it. A draft addressing identity chaining is in progress but not yet ratified. Document the scenario, record any federation workaround currently in use, and mark it as waiting on that draft.
Asynchronous long-running tasks. An agent that starts a task, goes dormant, then resumes hours later with the same token will hit expiry and scope-check failures that synchronous human flows never produce. Log the maximum token lifetime your current provider supports alongside the duration of your longest agent task. If they do not align, that is the gap.
Recursive delegation. Agent A spawns agent B with a narrowed scope. Current drafts place this pattern out of scope. If your product already runs this in production, document the trust chain and flag it as a future migration target.
Multi-user delegation. An agent enforcing delegated permissions on behalf of multiple human users simultaneously is a scenario where OAuth 2.1 may fall short. If this matches any current agent use case, it is the highest-priority item on your automation testing strategy backlog.
Naming these known unknowns now makes the migration scope visible before the standards force your hand.
What you have at the end of this process
Run through those steps and you have four concrete assets.
A named, layered gap log: every incompatibility between your current auth flow and the emerging agent-identity model is documented, prioritised, and tied to either RFC 9728 or a specific draft. Nothing is vague. Each entry has a layer, a current behaviour, an expected behaviour, and a blast-radius estimate.
An RFC 9728 compliance baseline: PKCE enforced, implicit grant removed, Protected Resource Metadata returning accurate scopes, sender-constrained refresh tokens in place. Covering it now removes the fastest route to a forced migration, which matters as the product life cycle compresses and experience gaps surface faster at every stage.
A reusable test actor: a scripted client that presents credentials programmatically, requests minimum scope, and retries on 401. Every future ai agent testing task your team runs starts from this asset, not from scratch.
A shared credential inventory: the shared credential inventory, with revocation tests confirming blast radius for each entry.
The standards will land. The teams that tested against the drafts will have a backlog to work through. The teams that did not will have a crisis.
Conclusion
Agent identity is not a future problem you can defer. The standards are forming now, the attack surface is real today, and the teams moving first will convert compliance work into a structured backlog rather than an emergency.
The four assets you build through this process, the gap log, the RFC 9728 baseline, the reusable test actor, and the shared credential inventory, are not throwaway audit artifacts. They become the foundation every future agent integration builds on.
Three things to take away: test against drafts early, document gaps without treating them as blockers, and name your shared credentials before a breach names them for you.
Start with Step 1 this week. Run your current auth flow against the MCP OAuth 2.1 mandate, log what breaks, and let the backlog take shape. Preparation done now is the crisis you avoid later.