e2e: TesterArmy's Agent Tests That Stop Calling the Model Once They Pass
The top repo on GitHub Trending on Sunday is a testing framework, and the clever part is when it doesn't use AI. e2e, from TesterArmy, is an Apache 2.0 end-to-end testing framework for web and mobile apps. It picked up 344 stars in a day and is near 3,000 total.
You write a test as a goal in plain English. agent.act("upgrade the workspace to the Pro plan"), then agent.assert("the invoice preview shows a prorated amount"), then an ordinary Playwright-style locator check in the same test. An agent drives the browser (Chromium, Firefox, WebKit through Playwright) or an iOS or Android simulator to reach the goal.
The part that matters is the cache. Once a later assertion has verified an agent step, its actions get recorded, and the next run replays them with no model calls until the app changes. So you pay for the model once per UI change, not once per CI run. That fixes the two complaints that have kept AI testing out of real pipelines: it's slow and costly, and it's flaky in a non-deterministic way. Tests without agent steps need no model at all. You can bring a subscription, an API key or a local model.
Two more details point at where the category is going. There's a decision-model executor package for "bounded semantic actions and assertions", which is the first place we've seen the decision-model wave land in a testing tool. And the whole docs site ships inside the npm package, so coding agents can read the docs offline from node_modules. That makes the framework itself built for agents to write tests with.
Link: github.com/tester-army/e2e and e2e.tester.army/docs
← Back to all articles
You write a test as a goal in plain English. agent.act("upgrade the workspace to the Pro plan"), then agent.assert("the invoice preview shows a prorated amount"), then an ordinary Playwright-style locator check in the same test. An agent drives the browser (Chromium, Firefox, WebKit through Playwright) or an iOS or Android simulator to reach the goal.
The part that matters is the cache. Once a later assertion has verified an agent step, its actions get recorded, and the next run replays them with no model calls until the app changes. So you pay for the model once per UI change, not once per CI run. That fixes the two complaints that have kept AI testing out of real pipelines: it's slow and costly, and it's flaky in a non-deterministic way. Tests without agent steps need no model at all. You can bring a subscription, an API key or a local model.
Two more details point at where the category is going. There's a decision-model executor package for "bounded semantic actions and assertions", which is the first place we've seen the decision-model wave land in a testing tool. And the whole docs site ships inside the npm package, so coding agents can read the docs offline from node_modules. That makes the framework itself built for agents to write tests with.
Link: github.com/tester-army/e2e and e2e.tester.army/docs
Comments