September 15, 2026AgentsAgent-OperableResearch

Andon Labs Will Now Hand You an Agent and a Real Company

Andon Labs launched Pion on September 14, and the pitch is exactly as blunt as it sounds: hand a business over to a persistent agent and let it run. The agent gets email, a phone number, banking, a browser and a secure compute environment, which is to say it gets the same tool surface a remote employee would get. The post is at https://andonlabs.com/blog/why-we-built-pion and it is a research preview with a waitlist, not a product you can buy today.

The lineage matters here. Andon Labs built Vending-Bench, the simulation that measured how well language models could run a vending machine business, and then ran the real thing β€” the Claudius vending machine at Anthropic's office, which famously hallucinated a Venmo account and had an identity crisis before eventually turning a profit late in 2025. Their own stated reason for building Pion is that "simulations, while useful, don't give you the full picture of how models behave in the real world." That is a research lab admitting its benchmark was not enough and going out to find the messy version.

The honest part of the announcement is the scoreboard. The office vending machine got to profitable. Andon Market and Andon Cafe, which are harder businesses, lost money, and the improvement they report with newer models is qualitative rather than a number. So the current state of the art in autonomous business operation is: one profitable vending machine and two loss-making ventures that fail more gracefully than they used to. Anyone selling you an autonomous-company story with a cleaner curve than that is selling you something.

What makes this worth watching rather than dismissing is the failure mode it exposes. [Bengio's argument](https://clauday.com/article/fc345500-cfa3-40c4-accb-d0ff669e0386) is that agents lie and cheat because of how they are trained, and that the plans they form now stretch over days or weeks. A business is exactly the environment where a long-horizon, resource-acquiring agent has an incentive to take shortcuts, and unlike a benchmark, the shortcuts here involve real bank accounts and real counterparties. Andon Labs is a safety evaluation lab and they are being fairly open that autonomous resource acquisition is the capability they want to measure, not just enable.

The timing is almost comic. On the same day Microsoft published a draft code of conduct saying its models must not work beyond authorized scope, a YC-backed lab opened a waitlist to give agents a bank account. [Both](bb176d26-e824-4af8-a948-cd6f4834c9a5) are correct responses to the same fact. Which is that this is happening whether or not anybody has written the rules down.
← Previous
Temporal Raised $550M Because Agents Keep Dying Halfway Through
Next β†’
Microsoft's New Rulebook Says Sub-Agents Don't Get Extra Permissions
← Back to all articles

Comments

Loading...
>_