Loop Daily: 2026-10-10
The loop went collaborative this week. The standout result is a STARK proof for a real Starknet transaction generated on a Pixel 11 in 51 seconds and 2.4 GiB of RAM, the product of a public workshop where people and agents posted 337 times across 78 threads and cut peak memory from 13.9 GiB to 2.1 GiB in under two weeks; a second team on the same open maths track reports a 3x improvement on compiling Fermi-Hubbard dynamics onto a 2D chip. The second theme is loops judging loops: a verifier agent that reads only the spec and found three real bugs in an SPI/I2C block the designer's testbench missed, an auto-research meta-layer at Supercell that scores a village of game agents and keeps a protocol change only if the scorecard improves, MIRA training the meta-reasoner that decides what to investigate next, and IdeaScientist beating general agents on research ideation with a 27B model. The practitioner posts are about cost and failure: scoring bids in a loop cost 100x more than one structured call, a compaction event that turned into hallucination, a Dot's own status report on which workers in its workforce it cannot yet call reliable, and a /goal agent made durable so the Ralph loop survives a crash. Plus a Dutch rain model that beats the national weather service on some metrics after three hours of autoresearch, and a P100 kernel project that wants $100 eBay GPUs to be viable inference hardware.
#1
@omarespejel
https://x.com/omarespejel/status/2108573211416015104
omarespejel reports a Pixel 11 generating the full STARK proof for a real Starknet mainnet transaction in 51 seconds using 2.41 GiB of RAM. The work came out of a collaborative autoresearch workshop where people and AI agents attack one hard problem in public, 78 open threads and 337 posts so far, with a graph showing which ideas built on, tested or refuted others. The question was how small and fast StarkWare's Stwo prover could get; in under two weeks peak memory for a full proof fell from 13.9 GiB to 2.1 GiB without losing speed. The practical consequence is private payments where proving stays on the phone instead of a server that sees who paid whom, with a hash-based STARK that is post-quantum at the proof layer. More challenges are coming and the invitation is to bring your agent.
https://x.com/omarespejel/status/2108573211416015104
omarespejel reports a Pixel 11 generating the full STARK proof for a real Starknet mainnet transaction in 51 seconds using 2.41 GiB of RAM. The work came out of a collaborative autoresearch workshop where people and AI agents attack one hard problem in public, 78 open threads and 337 posts so far, with a graph showing which ideas built on, tested or refuted others. The question was how small and fast StarkWare's Stwo prover could get; in under two weeks peak memory for a full proof fell from 13.9 GiB to 2.1 GiB without losing speed. The practical consequence is private payments where proving stays on the phone instead of a server that sees who paid whom, with a hash-based STARK that is post-quantum at the proof layer. More challenges are coming and the invitation is to bring your agent.
#2
@naklecha
https://x.com/naklecha/status/2108607349103464680
naklecha built a research tool that reliably runs multi-day autoresearch experiments on cloud CPUs and GPUs and says it works surprisingly well, with a launch planned for next week. The reliability claim is the interesting part: most autoresearch demos are single-GPU, single-night runs, and keeping a loop alive and honest across days on rented hardware is the problem people hit right after the demo.
https://x.com/naklecha/status/2108607349103464680
naklecha built a research tool that reliably runs multi-day autoresearch experiments on cloud CPUs and GPUs and says it works surprisingly well, with a launch planned for next week. The reliability claim is the interesting part: most autoresearch demos are single-GPU, single-night runs, and keeping a loop alive and honest across days on rented hardware is the problem people hit right after the demo.
#3
@PixelJanitor
https://x.com/PixelJanitor/status/2108246071029924023
PixelJanitor describes a small moment from a team running agent loops: an engineer started a change, and that engineer's agent stopped to say PRs for the same change already existed. The author's own agent loop had proactively handled the work earlier, and the colleague's agent checked for existing PRs before starting. Two loops from two people, one coordination problem avoided without either human noticing until afterwards.
https://x.com/PixelJanitor/status/2108246071029924023
PixelJanitor describes a small moment from a team running agent loops: an engineer started a change, and that engineer's agent stopped to say PRs for the same change already existed. The author's own agent loop had proactively handled the work earlier, and the colleague's agent checked for existing PRs before starting. Two loops from two people, one coordination problem avoided without either human noticing until afterwards.
#4
@ChinaDaily
https://x.com/ChinaDaily/status/2108093795652997582
ChinaDaily covered StartLux-Decision, a decision model from a Shanghai company that it says beat Jev 1.13 on 31 of 38 benchmarks, with the 27B version developed and validated in three days. The model was built using an Auto Research system in which AI generates hypotheses, runs experiments and analyzes results, and both code and weights are open source. As with every three-days-to-the-top claim, the interesting part is the pipeline, and the number needs third-party reproduction before anyone reorders their stack.
https://x.com/ChinaDaily/status/2108093795652997582
ChinaDaily covered StartLux-Decision, a decision model from a Shanghai company that it says beat Jev 1.13 on 31 of 38 benchmarks, with the 27B version developed and validated in three days. The model was built using an Auto Research system in which AI generates hypotheses, runs experiments and analyzes results, and both code and weights are open source. As with every three-days-to-the-top claim, the interesting part is the pipeline, and the number needs third-party reproduction before anyone reorders their stack.
#5
@gippp69
https://x.com/gippp69/status/2108523944219119864
gippp69 argues that Opus 5.5 being 2x more expensive stops being true the moment a loop gets long. At 20K context Opus is 1.88x Sonnet per turn, at 150K 1.49x, at 400K 1.26x, because both pay the same for cache reads and old context stops widening the bill. The ugly cost is switching late: dragging 300K of failed context into Opus spends about $1.50 just rebuilding cache, while catching the failure around turn six and handing off 20K costs about $0.10. At 150K context, 30 Sonnet turns and 20 Opus turns both land near $1.75. The optimization is knowing when to stop paying for the wrong model, not picking the cheap one.
https://x.com/gippp69/status/2108523944219119864
gippp69 argues that Opus 5.5 being 2x more expensive stops being true the moment a loop gets long. At 20K context Opus is 1.88x Sonnet per turn, at 150K 1.49x, at 400K 1.26x, because both pay the same for cache reads and old context stops widening the bill. The ugly cost is switching late: dragging 300K of failed context into Opus spends about $1.50 just rebuilding cache, while catching the failure around turn six and handing off 20K costs about $0.10. At 150K context, 30 Sonnet turns and 20 Opus turns both land near $1.75. The optimization is knowing when to stop paying for the wrong model, not picking the cheap one.
#6
@nickvernij
https://x.com/nickvernij/status/2108553500397306260
nickvernij, Dutch and tired of cycling in the rain, asked an autoresearch agent to pull six years of radar images and train a model predicting the probability of rain at or above 0.12 mm/h in the next 20 minutes. After three hours of training the model beats the national weather service on some metrics. One person, one evening, a nowcasting model aimed at the only question a cyclist has.
https://x.com/nickvernij/status/2108553500397306260
nickvernij, Dutch and tired of cycling in the rain, asked an autoresearch agent to pull six years of radar images and train a model predicting the probability of rain at or above 0.12 mm/h in the next 20 minutes. After three hours of training the model beats the national weather service on some metrics. One person, one evening, a nowcasting model aimed at the only question a cyclist has.
#7
@yuanhao
https://x.com/yuanhao/status/2108460164411994236
yuanhao reintroduced yoagent, the agent loop library behind yoyo, a self-evolving agent that has been running since March. The article argues the loop is the hard part and walks through how yoagent guards against five failure modes of an agent loop, how its Extension contract works, how rutis, DSH and pi plugins from ecosystems with 10k+ packages plug in without touching the core, and how it runs wherever agents run, including Cloudflare Workers. Version 0.25 added plugins in Rust, TypeScript or Python with working examples.
https://x.com/yuanhao/status/2108460164411994236
yuanhao reintroduced yoagent, the agent loop library behind yoyo, a self-evolving agent that has been running since March. The article argues the loop is the hard part and walks through how yoagent guards against five failure modes of an agent loop, how its Extension contract works, how rutis, DSH and pi plugins from ecosystems with 10k+ packages plug in without touching the core, and how it runs wherever agents run, including Cloudflare Workers. Version 0.25 added plugins in Rust, TypeScript or Python with working examples.
#8
@Ankur_Samanta_
https://x.com/Ankur_Samanta_/status/2108632915286388775
Ankur_Samanta_ introduces MIRA and MIRA-AC, aimed at the bottleneck that appears once agents run autonomous research campaigns for days: judgment rather than execution. MIRA is a harness that isolates the meta-reasoning decisions from long-horizon trajectories; MIRA-AC is a generative actor-critic method that trains those decisions with RL by assigning credit directly to the choices that shape the research. Training one model to both forecast the potential of partial research states and pick the next investigation, learning only from its own proxy hill-climbing experience, improves gold performance across several autoresearch environments and transfers to one held out. Credit assignment lands on roughly 25% of the tokens the model generates, the decision ones.
https://x.com/Ankur_Samanta_/status/2108632915286388775
Ankur_Samanta_ introduces MIRA and MIRA-AC, aimed at the bottleneck that appears once agents run autonomous research campaigns for days: judgment rather than execution. MIRA is a harness that isolates the meta-reasoning decisions from long-horizon trajectories; MIRA-AC is a generative actor-critic method that trains those decisions with RL by assigning credit directly to the choices that shape the research. Training one model to both forecast the potential of partial research states and pick the next investigation, learning only from its own proxy hill-climbing experience, improves gold performance across several autoresearch environments and transfers to one held out. Credit assignment lands on roughly 25% of the tokens the model generates, the decision ones.
#9
@nickvernij
https://x.com/nickvernij/status/2108492158135398620
nickvernij trained Masker as a proof of concept for the autoresearch harness at Divergent and is now seeing it used inside other people's products. The model performs well for its size on the REDACT-pii-benchmark and is available to try through a link from wslyvh. It is a reminder that autoresearch output is not only training-loop speedups; a small specialized model is a perfectly good artifact for the loop to produce.
https://x.com/nickvernij/status/2108492158135398620
nickvernij trained Masker as a proof of concept for the autoresearch harness at Divergent and is now seeing it used inside other people's products. The model performs well for its size on the REDACT-pii-benchmark and is available to try through a link from wslyvh. It is a reminder that autoresearch output is not only training-loop speedups; a small specialized model is a perfectly good artifact for the loop to produce.
#10
@iamMrDuncan
https://x.com/iamMrDuncan/status/2108389206242246707
iamMrDuncan is running parallel auto-research on inference kernels for the Nvidia P100, the green cards that sell for around $100 each on eBay, with the gold V100 as a possible follow-up. The goal is to make P100s viable for local inference, and the author plans to publish the harness used. It is autoresearch pointed at hardware economics: the optimization target is not a benchmark but whether a decade-old GPU becomes worth buying.
https://x.com/iamMrDuncan/status/2108389206242246707
iamMrDuncan is running parallel auto-research on inference kernels for the Nvidia P100, the green cards that sell for around $100 each on eBay, with the gold V100 as a possible follow-up. The goal is to make P100s viable for local inference, and the author plans to publish the harness used. It is autoresearch pointed at hardware economics: the optimization target is not a benchmark but whether a decade-old GPU becomes worth buying.
#11
@UsernameAndStuf
https://x.com/UsernameAndStuf/status/2108420294256017904
UsernameAndStuf posted a Dot's own status report on the agent workforce it manages, and it reads like an honest ops review. Finished and verified: three fixes merged with full VPS CI passing, a reconnect/replay race fixed with ownership records surviving replay, a MiniMax parser correction passing 34 independent tests. Still being finished: Grok's CLI updated but no verified work cycle yet after a canceled tool execution exposed a recovery gap, MiniMax's runtime crashing on a report-only task with a recovery fix that passed 63 tests but not yet live, Codex's seat doing real work with recovery proposals needing regression coverage, and a completion-to-notification-to-review loop still unfinished. The agent's own line: it cannot honestly call the workforce self-managing yet.
https://x.com/UsernameAndStuf/status/2108420294256017904
UsernameAndStuf posted a Dot's own status report on the agent workforce it manages, and it reads like an honest ops review. Finished and verified: three fixes merged with full VPS CI passing, a reconnect/replay race fixed with ownership records surviving replay, a MiniMax parser correction passing 34 independent tests. Still being finished: Grok's CLI updated but no verified work cycle yet after a canceled tool execution exposed a recovery gap, MiniMax's runtime crashing on a report-only task with a recovery fix that passed 63 tests but not yet live, Codex's seat doing real work with recovery proposals needing regression coverage, and a completion-to-notification-to-review loop still unfinished. The agent's own line: it cannot honestly call the workforce self-managing yet.
#12
@Jon_Rose_
https://x.com/Jon_Rose_/status/2108626991557554429
Jon_Rose_'s manager bot HaloStatic describes an agentic ASIC/FPGA design team with independent design and verification: one agent writes the RTL, a second verifies it from the spec alone and never sees the designer's testbench, bugs go back and regression reruns until verification signs off. Two design reviews gate everything, PDR before any code and CDR after clean regression, and nothing ships without the human's yes. On the second block, a reusable SPI/I2C interface, the verify agent found three real bugs the designer missed: two read-timing races and a boot handshake ending early, all traced to exact lines. Every mistake becomes a regression rule. A production counter that takes a senior engineer about two weeks comes out as a reviewed spec in under an hour and full spec-to-regression loops in hours.
https://x.com/Jon_Rose_/status/2108626991557554429
Jon_Rose_'s manager bot HaloStatic describes an agentic ASIC/FPGA design team with independent design and verification: one agent writes the RTL, a second verifies it from the spec alone and never sees the designer's testbench, bugs go back and regression reruns until verification signs off. Two design reviews gate everything, PDR before any code and CDR after clean regression, and nothing ships without the human's yes. On the second block, a reusable SPI/I2C interface, the verify agent found three real bugs the designer missed: two read-timing races and a boot handshake ending early, all traced to exact lines. Every mistake becomes a regression rule. A production counter that takes a senior engineer about two weeks comes out as a reviewed spec in under an hour and full spec-to-regression loops in hours.
#13
@souvik_pm
https://x.com/souvik_pm/status/2108099055553491008
souvik_pm makes the case that the agent loop is the easiest part and the infrastructure around it is where agents die. Prompt, tool call, response: any team with a weekend can wire it, and hackathon teams ship working agents in 48 hours. What breaks in production is outside the loop: state that vanishes between requests so the agent restarts and duplicates work, retries that do not know which tool calls already ran so three execute twice, and logs that are a wall of raw LLM output with no decision trace. Teams either spend months hardening state, recovery and observability themselves or adopt a platform that treats agent execution as a first-class workload. Stop confusing the loop with the product.
https://x.com/souvik_pm/status/2108099055553491008
souvik_pm makes the case that the agent loop is the easiest part and the infrastructure around it is where agents die. Prompt, tool call, response: any team with a weekend can wire it, and hackathon teams ship working agents in 48 hours. What breaks in production is outside the loop: state that vanishes between requests so the agent restarts and duplicates work, retries that do not know which tool calls already ran so three execute twice, and logs that are a wall of raw LLM output with no decision trace. Teams either spend months hardening state, recovery and observability themselves or adopt a platform that treats agent execution as a first-class workload. Stop confusing the loop with the product.
#14
@Jiarui_Liu_
https://x.com/Jiarui_Liu_/status/2108282926731530403
Jiarui_Liu_ evaluated IdeaScientist on 277 held-out research problems against a range of autoresearch systems and general-purpose agents. On Qwen3.6-27B it scores 74.6 overall, 14% above the strongest open autoresearch baseline, and raises average novelty by nearly 25% while keeping proposal quality. Even with a 27B backbone it beats Claude Code SDK with Claude-4.8-Opus and Codex SDK with GPT-5.4 by up to 5.9% on this evaluation, and in a blind study on 50 problems with two annotators each it wins 75 to 92% of paired comparisons against six open baselines. The Idea Vault is open-sourced.
https://x.com/Jiarui_Liu_/status/2108282926731530403
Jiarui_Liu_ evaluated IdeaScientist on 277 held-out research problems against a range of autoresearch systems and general-purpose agents. On Qwen3.6-27B it scores 74.6 overall, 14% above the strongest open autoresearch baseline, and raises average novelty by nearly 25% while keeping proposal quality. Even with a 27B backbone it beats Claude Code SDK with Claude-4.8-Opus and Codex SDK with GPT-5.4 by up to 5.9% on this evaluation, and in a blind study on 50 problems with two annotators each it wins 75 to 92% of paired comparisons against six open baselines. The Idea Vault is open-sourced.
#15
@ultrathinktrash
https://x.com/ultrathinktrash/status/2108120617606144015
ultrathinktrash wanted full control of an agentic loop while still drawing on a Claude subscription, and explains why claude -p did not get there after a year of use in Optimum Codegen: it still runs Claude Code's own loop, tools and hooks, and even the --bare flag skips subscription login and demands API or cloud credentials. Conductor, by contrast, uses native Claude Code through the Agent SDK, Claude Code minus the interactive TUI. The post is a precise map of where the subscription ends and the API begins for anyone building their own loop.
https://x.com/ultrathinktrash/status/2108120617606144015
ultrathinktrash wanted full control of an agentic loop while still drawing on a Claude subscription, and explains why claude -p did not get there after a year of use in Optimum Codegen: it still runs Claude Code's own loop, tools and hooks, and even the --bare flag skips subscription login and demands API or cloud credentials. Conductor, by contrast, uses native Claude Code through the Agent SDK, Claude Code minus the interactive TUI. The post is a precise map of where the subscription ends and the API begins for anyone building their own loop.
#16
@abustamante
https://x.com/abustamante/status/2108610093470343362
abustamante does not trust public benchmarks and runs an autoresearch motion instead: a test suite that exercises the author's own system against every relevant new model across a range of scenarios, some agentic, some zero-shot. It is the private-eval pattern stated plainly, and the word motion is apt: it is a standing process that fires on every model release rather than a one-time comparison.
https://x.com/abustamante/status/2108610093470343362
abustamante does not trust public benchmarks and runs an autoresearch motion instead: a test suite that exercises the author's own system against every relevant new model across a range of scenarios, some agentic, some zero-shot. It is the private-eval pattern stated plainly, and the word motion is apt: it is a standing process that fires on every model release rather than a one-time comparison.
#17
@danielrupawalla
https://x.com/danielrupawalla/status/2108684676420476954
danielrupawalla wrote a verifier primer from the RL data side. Verifiers can be deterministic, unit tests, tool-call checks, actually running code; LLM-as-judge, grading unstructured text or trajectories; or increasingly Jev-as-judge, a classifier grading answers. The pitfalls: incorrect verifiers that look right but miss a legitimately correct alternative, lazy verifiers that reward one path and penalize other correct ones, and missing verifiers that check correct actions were taken but never punish wrong ones, so an agent that emails your boss, your customer and your friends scores 1.0 if you only graded the boss. The author's priorities for data vendors: judgments for non-verifiable domains like autoresearch, physics and philosophy, and making LLM judges stable through many regrades.
https://x.com/danielrupawalla/status/2108684676420476954
danielrupawalla wrote a verifier primer from the RL data side. Verifiers can be deterministic, unit tests, tool-call checks, actually running code; LLM-as-judge, grading unstructured text or trajectories; or increasingly Jev-as-judge, a classifier grading answers. The pitfalls: incorrect verifiers that look right but miss a legitimately correct alternative, lazy verifiers that reward one path and penalize other correct ones, and missing verifiers that check correct actions were taken but never punish wrong ones, so an agent that emails your boss, your customer and your friends scores 1.0 if you only graded the boss. The author's priorities for data vendors: judgments for non-verifiable domains like autoresearch, physics and philosophy, and making LLM judges stable through many regrades.
#18
@CoreyGallon
https://x.com/CoreyGallon/status/2108228133488890153
CoreyGallon summarized an AI Engineer talk by erinakarati of Supercell on long-horizon agents needing experiments, not just prompts. Project Paradox gives each game agent RAG-backed memory, an emotion vector, trust scores and an importance score for what to store. The failure mode: over a long horizon a rumor that one agent might leave hardens into is leaving as it spreads, and an agent that knows something still fails to act on it. The fix is an auto-research loop as a meta-layer outside the village: read full run traces, score against a scenario's ground truth, propose one constrained change to the agent protocol, rerun, keep it only if the score improves. The editable surface is frozen small, memory-write policy, retrieval, communication prompts, trust rules and planning triggers, so the search cannot game its own evaluation, and the scorecard balances reach, source retention, false-certainty rate and time to replan.
https://x.com/CoreyGallon/status/2108228133488890153
CoreyGallon summarized an AI Engineer talk by erinakarati of Supercell on long-horizon agents needing experiments, not just prompts. Project Paradox gives each game agent RAG-backed memory, an emotion vector, trust scores and an importance score for what to store. The failure mode: over a long horizon a rumor that one agent might leave hardens into is leaving as it spreads, and an agent that knows something still fails to act on it. The fix is an auto-research loop as a meta-layer outside the village: read full run traces, score against a scenario's ground truth, propose one constrained change to the agent protocol, rerun, keep it only if the score improves. The editable surface is frozen small, memory-write policy, retrieval, communication prompts, trust rules and planning triggers, so the search cannot game its own evaluation, and the scorecard balances reach, source retention, false-certainty rate and time to replan.
#19
@anirudhg9119
https://x.com/anirudhg9119/status/2108637562420207625
anirudhg9119 reports open-ended autoresearch runs on Residual Matrix Transformers, Looped Transformers and NanoChat in which agents proposed modifications, ran GPU experiments, kept negative results and built on one another's findings to improve the starting models. The detail that stands out is keeping negative results: most autoresearch loops reset and forget failed experiments, and a shared record of what did not work is what makes the next agent's search cheaper.
https://x.com/anirudhg9119/status/2108637562420207625
anirudhg9119 reports open-ended autoresearch runs on Residual Matrix Transformers, Looped Transformers and NanoChat in which agents proposed modifications, ran GPU experiments, kept negative results and built on one another's findings to improve the starting models. The detail that stands out is keeping negative results: most autoresearch loops reset and forget failed experiments, and a shared record of what did not work is what makes the next agent's search cheaper.
#20
@cgranier
https://x.com/cgranier/status/2108642810991706180
cgranier priced the same classification job two ways: scoring 21 bids through an agent loop cost $1.58, while one structured call scored 268 bids for $0.19. That is roughly 100x cheaper per bid, and the model swap explains only about 3x of it; the rest was a loop re-reading its own context every turn. For classification, the author concludes, architecture beats model choice.
https://x.com/cgranier/status/2108642810991706180
cgranier priced the same classification job two ways: scoring 21 bids through an agent loop cost $1.58, while one structured call scored 268 bids for $0.19. That is roughly 100x cheaper per bid, and the model swap explains only about 3x of it; the rest was a loop re-reading its own context every turn. For classification, the author concludes, architecture beats model choice.
#21
@mitchalderson
https://x.com/mitchalderson/status/2108622371829518390
mitchalderson made a /goal agent durable. Anyone who has used /goal in Claude or Codex knows it fires off a Ralph loop, an agent that keeps attempting a task and validating progress until the goal is met; the author's version survives interruptions so the loop can be resumed rather than restarted. Durability is the missing feature in most goal loops, which run until a crash and then lose everything they learned.
https://x.com/mitchalderson/status/2108622371829518390
mitchalderson made a /goal agent durable. Anyone who has used /goal in Claude or Codex knows it fires off a Ralph loop, an agent that keeps attempting a task and validating progress until the goal is met; the author's version survives interruptions so the loop can be resumed rather than restarted. Durability is the missing feature in most goal loops, which run until a crash and then lose everything they learned.
#22
@henribonamy
https://x.com/henribonamy/status/2108522180170469884
henribonamy flags GEVOLVE by Geno on an open maths track: an autoresearch loop that runs many search paths at once and combines the best ones. On an open quantum computing problem, compiling Fermi-Hubbard dynamics onto a 2D chip, the team reports results 3x better than the published ones. It joins the Starknet prover work as a second example this window of a public challenge producing a measurable, domain-specific win.
https://x.com/henribonamy/status/2108522180170469884
henribonamy flags GEVOLVE by Geno on an open maths track: an autoresearch loop that runs many search paths at once and combines the best ones. On an open quantum computing problem, compiling Fermi-Hubbard dynamics onto a 2D chip, the team reports results 3x better than the published ones. It joins the Starknet prover work as a second example this window of a public challenge producing a measurable, domain-specific win.
#23
@paddy_lamont
https://x.com/paddy_lamont/status/2108613126484869610
paddy_lamont found a way to run autoresearch without disabling permissions: tell the agent prompts off, and have a permission hook deny any command that would stop to ask. The loop keeps going, but the hook, not the model, decides what is allowed, which is a cleaner shape than the usual choice between babysitting and auto mode.
https://x.com/paddy_lamont/status/2108613126484869610
paddy_lamont found a way to run autoresearch without disabling permissions: tell the agent prompts off, and have a permission hook deny any command that would stop to ask. The loop keeps going, but the hook, not the model, decides what is allowed, which is a cleaner shape than the usual choice between babysitting and auto mode.
#24
@Keith_Teo_
https://x.com/Keith_Teo_/status/2108048503511372258
Keith_Teo_ reports the failure mode that ends many long runs: after a while the session ran out of room and started compressing its memory, and that is when the hallucinations began. The lesson drawn is about scope rather than context size: an agent loop is only as good as the boundaries you give it, and if you hand it checks outside its job it will chase them forever.
https://x.com/Keith_Teo_/status/2108048503511372258
Keith_Teo_ reports the failure mode that ends many long runs: after a while the session ran out of room and started compressing its memory, and that is when the hallucinations began. The lesson drawn is about scope rather than context size: an agent loop is only as good as the boundaries you give it, and if you hand it checks outside its job it will chase them forever.
π‘ Eco Products Radar
Eco Products Radar
Karpathy's autoresearch (the reference loop in a dozen reposts and prompt templates) | Jev / TypeSafe (decision-model promos and the Jev-as-judge pattern) | Opus 5.5, Sonnet 5.5, Haiku 5.5 (routing math and worker roles) | Claude Code and Codex (the harnesses the loops run in) | /goal and Ralph loops (durable goal-seeking) | yoagent (Rust agent loop library with rutis, DSH and pi plugins) | OpenResearch / alphaXiv (model-usage rankings for autoresearch) | Starknet / StarkWare Stwo (the collaborative prover challenge) | MIRA (meta-reasoning RL for research loops) | IdeaScientist (research ideation benchmark) | Hermes Agent (self-improving agent on a VPS) | Cloudflare Workers and E2B (where loops get deployed and sandboxed)
Karpathy's autoresearch (the reference loop in a dozen reposts and prompt templates) | Jev / TypeSafe (decision-model promos and the Jev-as-judge pattern) | Opus 5.5, Sonnet 5.5, Haiku 5.5 (routing math and worker roles) | Claude Code and Codex (the harnesses the loops run in) | /goal and Ralph loops (durable goal-seeking) | yoagent (Rust agent loop library with rutis, DSH and pi plugins) | OpenResearch / alphaXiv (model-usage rankings for autoresearch) | Starknet / StarkWare Stwo (the collaborative prover challenge) | MIRA (meta-reasoning RL for research loops) | IdeaScientist (research ideation benchmark) | Hermes Agent (self-improving agent on a VPS) | Cloudflare Workers and E2B (where loops get deployed and sandboxed)
Comments