E2E Tests with AI That Don't Use Tokens: e2e Records the Agent and Replays without Calling the Model

e2e is an open source end-to-end (E2E) testing framework with AI: you describe an objective in natural language, an agent navigates your web or mobile app, and once the result is verified, its actions are repeated without calling the model until the interface changes.

That mechanism is the entire proposal. Most AI testing tools pay for inference on every run. e2e pays once, saves what worked, and only goes back to the model when the screen no longer matches the recording. It’s published by TesterArmy under the Apache-2.0 license.

What is e2e and how do its end-to-end tests with AI work?

e2e is a test runner, an SDK, and a CLI, published on npm as e2e. For web it controls Chromium, Firefox, and WebKit through Playwright; for mobile, iOS simulators and Android emulators. In a single test you can combine three types of steps:

  • Objectives for the agent: agent.act('upgrade the workspace to the Pro plan')
  • Verifications judged by AI: agent.assert('the invoice preview shows a prorated amount')
  • Traditional deterministic assertions: expect(screen.getByRole('status')).toContainText('Pro')

A test with no agent steps doesn’t need any model. That matters: you can adopt e2e as a conventional e2e testing runner and add AI only to flows where hand-written selectors break over and over.

How does the replay cache save model calls?

The cache records the actions of an agent.act() step only after a subsequent verification passes: an assertion on a locator, a URL check, or an agent.assert. On the next run, the runner:

  1. Checks that the starting screen is the same route.
  2. Finds each recorded control (by role, name, test id, and context) and repeats its action.
  3. Verifies the final route and the controls that appeared, disappeared, or changed state.
  4. If everything matches, it finishes the step without calling the model. If anything doesn’t match, the agent picks up from the current screen.

The summary of each run shows how many steps were replayed, how many were handed off to the agent midway, and how many were missed entirely.

Details worth knowing before counting on the savings:

  • Verifications aren’t cached. agent.assert, agent.waitFor, and agent.extract always run live. The savings are in action steps, not in judgments: a test full of AI assertions still calls the model on every run.
  • Dynamic values need unique(). A timestamp or a new email would cause a cache miss on every run; wrapping them in unique() makes the runner substitute the current value when replaying.
  • In CI the cache is read-only by default. Locally it reads and writes; in CI it only replays, unless you set cache: 'read-write'.
  • The cache isn’t shared by default. e2e init adds .e2e/cache/ to .gitignore. To share recordings with CI, remove that line, commit the directory, and review the entries like you’d review test code.
  • Stale recordings can fail visibly. Without extra configuration, if a recording stops working the agent takes the step silently and every CI run pays for model calls again. With --strict-cache, that becomes a REPLAY_STALE failure.

How do you install e2e?

In your app’s directory:

npx e2e init
npx e2e run

The init assistant asks you for an engine (web or mobile) and a model provider, or None if you want tests without AI, and writes e2e.config.ts and an example test. The first run downloads a browser or starts a simulator, checks that the app opens, and writes .e2e/report.json. It doesn’t make model calls, so you still don’t need an API key.

Requirements: Node.js 24.8 or higher (22.22.3 or higher on the 22 line). Mobile testing also requires Xcode with an iOS simulator runtime, or the Android SDK with an emulator.

Does e2e work on Windows?

Not natively: the documentation says to run it inside WSL.

Can you use it from Claude Code, Codex, or Cursor?

Yes, and it’s probably the fastest way to get started. The getting started guide includes a prompt you paste into Claude Code, Codex, Cursor, or another code agent: the agent runs init, installs the e2e skill, points the config to your dev server, asks you which model to use, and iterates until the example test passes.

After that, requests like “add a test for checkout” or “why did this run fail?” work without further setup. e2e also includes an MCP server (e2e mcp) so your code agent can inspect the app and handle it live while writing the test, and the package brings all the documentation in node_modules/e2e/docs so the agent can read it offline.

Is e2e free and what models can it use?

The framework is free and open source; what you pay for, if anything, is model usage in the agent steps. There are three paths:

Subscriptions you already have, no API key:

Subscription Command
ChatGPT Plus or Pro npx e2e login openai
GitHub Copilot npx e2e login github-copilot
OpenCode Console npx e2e login opencode-console
SuperGrok or X Premium+ npx e2e login spacexai

API keys from any Vercel AI SDK provider: Anthropic, OpenAI, Google, Mistral, DeepSeek, and OpenRouter, among many others. Claude is available this way, with an API key, not with subscription login.

Local models with Ollama, LM Studio, or any OpenAI-compatible endpoint, like vLLM or llama-server.

One tip: let the init assistant write the model line instead of copying a model ID from an article. IDs change faster than documentation.

e2e, Playwright, Cypress, or Selenium?

You don’t have to choose right away. @e2e-dev/web brings its own pinned playwright-core, so your @playwright/test is untouched and both runners coexist in the same project; e2e looks for tests/**/*.e2e.ts by default, so file patterns don’t clash. The documentation includes migration guides from Playwright, Cypress, Selenium, Detox, and Maestro.

The Playwright guide shows the tradeoff well. A checkout test with seven scripted lines (navigate, click Checkout, fill card, expiry and CVC, click Pay, final assertion) becomes this in e2e:

test('checkout', async ({ app, agent, screen }) => {
  await app.open('/cart');
  await agent.act('pay with the test card 4242 4242 4242 4242, expiry 12/30, CVC 123');
  await expect(screen.getByRole('heading')).toHaveText('Order confirmed');
});

Navigation and the final assertion stay deterministic; the agent solves the middle form and adapts if it changes. The guide itself recommends starting with flows whose steps change often, like wizards and checkouts, and leaving Playwright tests alone where they’re already stable.

It’s also honest about what’s missing: no pixel snapshot comparison yet, tests in the same file don’t run in parallel, and there’s no HTML reporter. Also, text locators search for exact match and are case-sensitive by design, whereas Playwright searches for substrings.

What should you know before adopting it?

  • It’s pre-1.0. At the time this post was published (October 2026), the README says the project is under active development toward 1.0 and APIs and config can still change between minor versions. Pin the version.
  • Secrets stay out of the model’s view. Passwords are passed as a Secret handle that the model never sees, and after filling a secret, screenshots are disabled for the rest of the test.
  • Test code doesn’t run in a sandbox. Tests and config run with your operating system’s permissions. The security docs recommend running code from untrusted pull requests in an external sandbox, without secrets.
  • Telemetry is on by default. The CLI sends anonymous usage data. To turn it off: npx e2e telemetry disable or E2E_TELEMETRY_DISABLED=1.
  • There’s a company behind it. TesterArmy sells a hosted agentic testing platform. The framework is fully Apache-2.0, but the README links point to the commercial product.

Is it worth trying?

If you have an end-to-end suite where every redesign breaks a dozen selectors per sprint, yes. The core idea is right: use the model for the part that changes and stop paying for the part that doesn’t. Start with a single fragile flow, commit the cache, run CI with --strict-cache, and watch the count of repeated steps go up.


What about you? Which flow in your app breaks your end-to-end tests every time the interface changes?

How do you do your end-to-end testing today?
  • Playwright
  • Cypress
  • Selenium
  • I prefer another option (tell us which)
0 votantes

Comment below, or if this article reached you by email, reply directly to the email: your reply gets posted here.

Related: What’s Changing with GitHub Copilot Browser Tools