All posts
PPremkumar Unnikrishnan··6 min read

Inside autoreg.md: The One File Your AI Agent Needs to Test Your App

Every test AutoReg generates traces back to a single file: autoreg.md. No dashboards to configure, no test framework to learn, no selectors to hand-maintain — just one markdown file the AI agent reads before it writes a single test. This post is about what's actually in that file, why you never write it by hand, and why that matters more than it sounds.

It's a spec, not a to-do list

Think of autoreg.md as a test brief written for a QA engineer who has never seen your app and never will get the chance to ask you a follow-up question mid-test-run. Everything the agent knows — your base URL, which routes need a login, the exact selector for the submit button, what the error toast says when a password is wrong — comes from this file plus whatever it observes live in your running app. If it's not in the spec and not on the page, the agent doesn't know it.

That's a deliberate design choice. Vague specs produce vague tests that fail for the wrong reasons. A precise one produces a suite that passes on the first run, because every claim in it was checked against something real before it was written down.

You don't write it — a skill does

Here's the part that surprises people: autoreg.md is never hand-written. It's produced by the generate-autoreg skill, which installs into Claude Code, Cursor, Copilot, Windsurf, and a dozen other AI coding agents in seconds. You point it at your app's source and run one slash command:

terminal
$ /generate-autoreg /path/to/workflow /path/to/app/src

The skill reads your actual source — routes, forms, validation schemas, seed data, the works — and only stops to ask you things it genuinely can't infer: test credentials for each role, whether writes are safe to perform, realistic sample values for domain-specific forms. It batches those questions into one message instead of drip-feeding them, and it never invents a password or account for a real backend.

Every selector that ends up in the spec is ranked by how likely it is to survive a redesign — data-testid and stable ids first, hashed CSS-module classes and auto-generated framework ids never. That ranking is the difference between a suite that still runs after your next UI tweak and one that needs repairing every sprint.

What's actually inside

A generated spec covers a lot more ground than “test the login page”:

  • Application context — what the app does, its base URL, tech stack, and rendering model, so the agent isn't guessing at basics.
  • Auth & test accounts — one row per role, the exact login flow, and how a session ends.
  • Pages and routes — every path in the app, including which ones require login, with the component or handler behind each.
  • Confirmed UI selectors — grouped by page, best selector first, every one of them checked against your live DOM.
  • Domain types and validation rules — copied verbatim from your TypeScript interfaces and Zod/Yup schemas, so negative test cases assert the exact error text your app actually shows.
  • Page behaviors — which interactions trigger a network call and a loading spinner, and which are instant client-side state changes that need no wait at all.
  • A test data catalog — every value any scenario fills into a form, tagged by where it came from: your answers, your seed files, or realistic synthetic data.
  • Invalid selectors and open questions — selectors an LLM might guess but that don't exist in your app, and a running list of anything still unconfirmed.

Nothing in that list is filler. Every section exists because a real failure mode needed it — the test data catalog exists so two scenarios never disagree on what “a valid email” means; the invalid-selectors list exists because a plausible-looking guess is worse than an obvious one.

Signed, so nobody has to take it on faith

Because the spec drives everything downstream, AutoReg needs to know a given autoreg.md is genuinely what the skill produced and hasn't been quietly edited since. So the last thing the skill does — automatically, with no manual step — is send the finished file for a cryptographic signature, appended as a footer at the very end of the file.

When you import the spec, AutoReg verifies that signature. A genuine, unaltered file imports clean. One that's been hand-edited after signing is rejected as tampered, and one that was never signed at all — say, generated offline — imports with an honest “unverified” warning instead of silently pretending to be trustworthy.

Better spec in, better suite out

Once autoreg.md is imported, AutoReg's agent reads it, enumerates every scenario the spec describes or implies, and generates one tests/scenario-NNN.json file per scenario — each one re-confirmed against a live DOM snapshot before it's written. The spec is the one input; the suite is the output. That's the whole reason a detailed, precise spec is worth the extra few minutes the skill spends interviewing you up front — it's the only lever you have over how good the generated tests are.

If your app changes, re-run the skill — don't edit

The one rule worth remembering: when your app changes, re-run generate-autoreg rather than hand-editing the file. Editing it by hand breaks the signature and means AutoReg can no longer vouch for anything the file claims.

Re-running is cheap, because it isn't a fresh start. The first run saves a baseline of your app — its features, routes, selectors and which scenario covers what. Every run after that diffs your source against it and regenerates only the scenarios your changes actually affect, tracing dependencies so a change to a shared component or an auth flow also re-verifies whatever sits downstream of it. Scenarios nothing touched are carried through untouched, so tests that already pass keep passing. You also get an autoreg.delta.md showing exactly what changed and why — and in AutoReg, asking the agent to “apply the spec delta” previews every add, update and delete before it touches a test.

Try it

The AutoReg: Skill Installer extension installs the generate-autoreg skill into 16+ coding agents — Claude Code, Cursor, Copilot, Windsurf, and more — in seconds. Point it at your app and see what a fully-sourced, signed autoreg.md looks like for your own codebase.

Usage