Introducing the Stunt Double Index
A free, public benchmark of how AI agents experience websites. Look up your domain, see where agents stop, and get the list of what to fix.
Today we're announcing the Stunt Double Index: a free, public benchmark of how AI agents experience websites. It currently covers more than 650 sites, and you can submit your own in one-click.
What you get
A comprehensive scorecard and report with:
- A score out of 100, with a breakdown across eight questions, so you can see which part of the experience is holding the number down.
- The frictions behind it: the specific places an agent was turned away, got lost, or could not finish. A bot challenge on the homepage. A pricing page that renders as an empty shell. A checkout with nowhere to start without signing in.
- A way to fix them. Each scorecard has a fix prompt you can hand to your own coding agent, and the same data is available over our MCP server if you would rather ask for it from Claude or Cursor.
How a site is scored
Full details on the methodology page.
Quick checks. Before any agent visits, we make a handful of requests to the site. Does robots.txt let each agent in? Is there a bot challenge at the door? Is the site monetised and can an agent find pricing, signup, cart and contact? Does the site publish what agents look for: llms.txt, structured data, an MCP server card, OAuth discovery? Each check keeps its evidence, so the scorecard shows what passed, what failed and why.
Agents trying real tasks. Claude, ChatGPT and Gemini open the site in a real browser and try what a customer would: research the company, then start at the homepage and then completing a task expected of the audience. We then benchmark performance with a separate grader against a fixed rubric, so no agent marks its own work.
The score is blended with agent tasks weighted for 60% and the quick checks for 40%.
| Question | Category | Weight |
|---|---|---|
| Can an agent complete a task on behalf of a user? | Task completion | 20% |
| Will they pick you? | Discovery | 15% |
| Can they read your site? | Information retrieval | 15% |
| Do they tell the truth about you? | Accuracy | 15% |
| Can agents recognise you exist? | Brand awareness | 10% |
| Where do you sit in the lineup? | Market ranking | 10% |
| Do you let agents in, safely? | Delegated access | 10% |
| Can an agent reach a human? | Contact & communication | 5% |
Limitations
Agent sessions cost real compute, so we're limiting the first batch of scorecards. The top 50 sites in the rankings are re-scored with full agent sessions every month and other sites use just the quick checks.
Some sites turn our requests away or do not answer at all; we mark those not scored and keep them out of the rankings. If a score looks wrong, you can help us fix it here.
How our agents behave on your site
Index is built with privacy, trust, and security in mind. Our agents behave predictably and respect your site's settings:
- They say who they are. Every request carries
StuntDoubleIndexin its user agent, so you can see us in your logs and name us inrobots.txt. - They respect your opt-out. If
robots.txtor your headers say no, we stay out. - They never pay or sign up for real. No payment goes through and no real registration details are entered.
- They keep no personal information. Sessions record what the agent saw and did.
- They visit gently. We cap how often we request pages from each site.
Getting a deeper look at your site
Claim your domain and you can watch the session replays, see every step the agents took, and request a full rescan after you ship a fix. If the score is one you are proud of, each scorecard has a badge you can embed that updates when the score does. You'll need to use a business email that matches the domain of the site you're claiming (@gmail, @yahoo, and other public email services are not allowed to claim domains).
Why we built it
Stunt Double puts AI actors through your own product and shows you where they get stuck. The Index is the same idea pointed at the public web, with the results open to everyone.
Performance and accessibility both became things teams fixed once there was a number anyone could look up. Agent experience has no such number. This is my attempt at one, with the method published so you can argue with it.
Check your site. If our agent got stuck there, an agent working for a real customer would too.
Keep reading
- 3 min
What Is MCP and Why It Matters for Product Teams
The Model Context Protocol lets AI assistants call tools in external services. For product teams, that means testing, feedback, and issue tracking without leaving your editor.
- 9 min
Your Users may not be Human
In June 2026 Cloudflare confirmed machines now generate more web traffic than people. Every one of those visits is a user nobody has watched, and nobody can ask why. What does user-centred mean in this new age of AI?
- 4 min
Building with Supabase
How we run a multi-tenant AI agent platform on Supabase: row level security for tenant isolation, OAuth and passkeys for humans and machines, pgvector for agent RAG, read replicas for global reads, and preview branches for every pull request.