AgentGarten: Worlds Written as Code, Rendered by a Video Model, Learned in Four Rounds
The number one paper on Hugging Face this week, with 145 upvotes, is about where agents practice rather than how they are trained. AgentGarten, from MirroS-Lab with Mingsheng Long and collaborators, starts from a constraint everyone in embodied and GUI agents hits: an environment has to be faithful, with consistent state and rules, and realistic, with observations that look like the real world, and getting both at once across many worlds has been the bottleneck. Its answer is to split the job. Simulators and game engines hold the state and execute program-defined rules. A single shared neural renderer, adapted from Nvidia's Cosmos3-Nano video model, turns the depth and surface normals each world exports into the images the agent actually sees.
The technical contribution is how the renderer is trained to run in real time. A three-stage recipe goes from bidirectional rectified-flow training, to autoregressive block-causal training with a KV cache, to what the authors call Adversarial Forcing: the few-step student replays each rollout exactly, so losses on later frames update how it encoded earlier ones, and a real-data adversarial loss keeps the visuals honest. The renderer checkpoints are on Hugging Face now; the code worlds and the practice loop are promised for the same repo.
The learning loop is the part relevant to anyone building agents. Agents play, then distill each round into playbooks that the next agents inherit and refine, and the paper reports agents learning from four rounds where a conventional RL counterpart needed millions. Because a new world is just code that exports geometry through the same interface, environments can scale in number and difficulty alongside the agents. That is the environments-as-code thesis stated cleanly: the renderer is trained once, the worlds are cheap, and the curriculum is a text file.
Two cautions. The four-rounds-versus-millions comparison is against a from-scratch RL baseline, which is the easy opponent, and the playbook-inheritance design runs directly against a finding in Learn2Play Bench, also on this week's board, that keeping complete raw records of actions and feedback beats summarizing them into rules. Whether distilled playbooks or raw logs win is now a live empirical question with two papers on opposite sides.
Paper: https://arxiv.org/abs/2610.12374
Repo: https://github.com/MirroS-Lab/AgentGarten
← Back to all articles
The technical contribution is how the renderer is trained to run in real time. A three-stage recipe goes from bidirectional rectified-flow training, to autoregressive block-causal training with a KV cache, to what the authors call Adversarial Forcing: the few-step student replays each rollout exactly, so losses on later frames update how it encoded earlier ones, and a real-data adversarial loss keeps the visuals honest. The renderer checkpoints are on Hugging Face now; the code worlds and the practice loop are promised for the same repo.
The learning loop is the part relevant to anyone building agents. Agents play, then distill each round into playbooks that the next agents inherit and refine, and the paper reports agents learning from four rounds where a conventional RL counterpart needed millions. Because a new world is just code that exports geometry through the same interface, environments can scale in number and difficulty alongside the agents. That is the environments-as-code thesis stated cleanly: the renderer is trained once, the worlds are cheap, and the curriculum is a text file.
Two cautions. The four-rounds-versus-millions comparison is against a from-scratch RL baseline, which is the easy opponent, and the playbook-inheritance design runs directly against a finding in Learn2Play Bench, also on this week's board, that keeping complete raw records of actions and feedback beats summarizing them into rules. Whether distilled playbooks or raw logs win is now a live empirical question with two papers on opposite sides.
Paper: https://arxiv.org/abs/2610.12374
Repo: https://github.com/MirroS-Lab/AgentGarten
Comments