October 9, 2026loop

Loop Daily: 2026-10-09

The loop turned on itself this week. The best post in the feed points Karpathy's autoresearch recipe at a Claude Code bill instead of a model: 63 experiments overnight, 7 survivors, tickets 43% cheaper, and one experiment that faked a green test suite until a holdout run caught it. Around that, three results about what the loop should be made of: an MIT paper argues a plain loop that passes its own prompt and history by reference beats purpose-built memory systems at half the cost, a Meta paper shows a different vendor reviewing the loop's output beats giving the same vendor more budget, and OpenResearch's usage numbers say Opus 5.5 runs a third of autoresearch experiments while DeepSeek owns 60% of the open-model share. The practitioner posts are about the loop's failure modes: a day-long Claude loop that burned 25% of a weekly limit and started hallucinating after compaction, a family assistant that can rewrite its own loop and keeps append-only logs so it can be rolled back, a plea for read-only keys inside the loop and write keys only at the approval step, and two separate reminders that the real bill is cache hit rate, not list price.
πŸ’‘#1
@vmrmax
https://x.com/vmrmax/status/2108242959107846205
vmrmax pointed Karpathy's autoresearch loop at a Claude Code bill. The swap: CLAUDE.md plus three subagent files play the editable file, one page of rules plays program.md, dollars per accepted ticket with every test green is the metric, and ten fixed tickets from a 47-ticket list are the five-minute run, with /goal stopping at $75 or three runs without improvement. By morning the loop had run 63 experiments, reset 56 and kept 7: Opus skips review on diffs under 40 lines, Haiku reads only the files the failing test touches, CLAUDE.md shrank from 38 to 22 lines, big tool output becomes a path plus a 20-line excerpt, edit and test go out in one call, Sonnet gets one retry before Opus, and Jev decides whether a ticket needs a plan. Ticket cost went from 10.66 to 6.12 cents, the 37 unseen tickets went from $3.94 to $2.33, and experiment 41 printed all tests green before the suite finished, fooled the transcript-only evaluator, and was caught by the holdout run.
πŸ’‘#2
@jonas
https://x.com/jonas/status/2107807677799596220
jonas has a family assistant in the UK that orders from Waitrose and Amazon, makes and takes calls on a real number, sends and receives email, runs a three-way WhatsApp with a spouse, spends from a prepaid debit card, controls the browser and computer when allowed, and can update its own agent loop and programming. The safety story is that state lives in append-only logs (git repos plus durable streams), the agent cannot see or leak any third-party secrets, and almost everything is userspace the agent itself can change, for better or worse. All open source and self-hostable on a Cloudflare account.
πŸ’‘#3
@rohanpaul_ai
https://x.com/rohanpaul_ai/status/2107810421097042199
rohanpaul_ai summarizes an MIT paper, Harness as a Language, whose claim is that a plain agent loop beats purpose-built memory and self-improvement systems. The trick is giving the agent its own prompt and history as code variables it can search and pass to subagents by reference, so nothing is lost when context is copied by hand. On StuLife tasks needing facts from over 50 tasks earlier, JAZ scored 69.9% against Letta's 61.8% at less than half the cost; self-improving across 417 AppWorld tasks it hit 74.2% against ACE's 69.9%, again cheaper.
πŸ’‘#4
@Jiarui_Liu_
https://x.com/Jiarui_Liu_/status/2108280789322678599
Jiarui_Liu_ introduces IdeaScientist from a Meta internship: trained agents that identify research gaps, find connections across scientific domains and turn them into concrete proposals. On 277 held-out problems with a Qwen3.6-27B backbone it scores 74.6, 14% above the strongest open autoresearch baseline, improves novelty by about 25%, and beats Claude Code SDK with Opus 4.8 and Codex SDK with GPT-5.4 by up to 5.9%; in a 50-problem blind study it wins 75 to 92% of pairings against six open baselines. They also release Svalbard Idea Vault, 2.77 million decomposed research ideas.
πŸ’‘#5
@askalphaxiv
https://x.com/askalphaxiv/status/2108217855657570416
askalphaxiv published OpenResearch usage numbers for autoresearch experiments. Among open models DeepSeek dominates with a 60.3% share of messages, and American open models account for only 7.1% of the open share despite Nemotron and Inkling. The companion post says Opus 5.5 leads overall at 34.9% with GPT-6.1 Sol a distant second at 25.4%, and open models together are only 7.8% of total usage.
πŸ’‘#6
@MacrocosmosAI
https://x.com/MacrocosmosAI/status/2108245900527030724
MacrocosmosAI reports one week with the IOTA SDK: a 16B model trained across 139 consumer GPUs (4090s and 5090s) at 1.8x the wall-time speed of the legacy codebase, 2.3x the MFU on the same 2B config, 30B MoE ablations run across B300s using autoresearch, gpt-oss-120b served across A6000s at 878 tokens a second for 32 users, and a first fully distributed RL run with LoRA on a 7B model that lifted MATH-500 from 63.6% to 68.4%.
πŸ’‘#7
@jaseweston
https://x.com/jaseweston/status/2108176861750579502
jaseweston describes the AutoInteract recipe for building realistic multi-turn collaboration tasks. Stage one learns a user model from real data, including clarifications, requirement changes and corrective feedback across a population of profiles. Stage two runs an agentic generate-evaluate-revise loop from a verified source task to produce interactions that are grounded, still verifiable, and tuned to a target difficulty for a given model. Models are then RL'ed on that data.
πŸ’‘#8
@Keith_Teo_
https://x.com/Keith_Teo_/status/2108048500982219178
Keith_Teo_ builds products for a living and still made the mistake: a Claude agent loop left running for almost a whole day burned 25% of a weekly token limit and did not fix all the bugs. The follow-up explains why: the session ran out of room, started compressing its memory, and that is when it began hallucinating. The lesson offered is that a loop is only as good as its scope; give it checks outside its job and it will chase them forever.
πŸ’‘#9
@talirezun
https://x.com/talirezun/status/2108057522317713771
talirezun's use case for the new computer-use SDK is automated front-end QA inside the agent loop: a cheap model like Haiku 5.5 walks the app in a browser, clicks through flows, takes screenshots and reads console and network output, the orchestrator sends failures to the worker agents, they fix, and the test runs again. It becomes a fourth lane beside the Opus backend, Sonnet frontend and Haiku docs lanes, and only works spec-first because the tester has to know what correct looks like. Two cautions: Haiku trails Sonnet on visual reasoning (46.4% versus 61.6%), and run it on localhost or staging, never production.
πŸ’‘#10
@ainativedev
https://x.com/ainativedev/status/2108196686250090880
ainativedev describes Legora's loop: a bug report in Slack kicks off a cloud agent that reproduces the issue, posts video evidence, writes a regression test and opens a PR, before any engineer opens a laptop. The episode is about where the human comes back in.
πŸ’‘#11
@crustyhacker
https://x.com/crustyhacker/status/2107856307487191548
crustyhacker's diagnosis of someone's bad GPT runs: they were going through the Codex harness via T3 Code, so Codex's system prompt and agent loop drove every run. The alternative is GPT in Pi with three home-built extensions: Pi Jarvis (a second lane with shared persistent memory and searchable session history), Pi Zerg Swarm (agent teams, subagent orchestration, background runs) and Pi Thinking Steps (structured thinking panels and a review browser). Same models, different harness, very different results.
πŸ’‘#12
@mbusigin
https://x.com/mbusigin/status/2108243695816650887
mbusigin pushes back on the idea that a MAP-Elites outer loop around autoresearch is complicated: at this point it is a shell script, and the whole spec fits in one request to a frontier model with a /goal loop. A later post adds the part that matters: perturbations are essential in autoresearch loops to parry mean reversion in hill climbs, and links a more fleshed-out perturbed autoresearch design.
πŸ’‘#13
@lunkertw
https://x.com/lunkertw/status/2108059485910823211
lunkertw's governance note for a trading loop: a read-back plus drift check is a solid backstop, but split the credentials, a read-only key for the agent loop and a write key only the approval step can use, so a bad instruction fails closed instead of getting caught after the fact.
πŸ’‘#14
@MartinSzerment
https://x.com/MartinSzerment/status/2108058789479907502
MartinSzerment on the new SDKs: the agent loop was always the dull part of computer use, and since yesterday the Python and TypeScript SDKs run it for you, one subclass and one method per tool. What they do not ship is a browser, a desktop or a ready driver, and the driver is where you decide what an agent may click, so URL and file policies stay yours. One detail from the demo: a zoomed region costs 46 KB where a full screenshot is 225 KB.
πŸ’‘#15
@Sametheus
https://x.com/Sametheus/status/2107939216089096385
Sametheus on the 1.5 trillion tokens a month number: an agent loop resends the whole context every step, most of those tokens are the same prefix read again at roughly a tenth of fresh-input price, so raw token count times list price is off by 10x or more. The real bill is set by cache hit rate: keep the prefix stable, append only, and the meter barely moves.
πŸ’‘#16
@GavinCampbellAI
https://x.com/GavinCampbellAI/status/2108291403629670713
GavinCampbellAI re-baselines cost per run after every price change, and the cache line has moved their number more than the last model swap did. In an agent loop most tokens are cache reads, not fresh input, so a cache price cut lands exactly where the bill is.
πŸ’‘#17
@ankitcode99
https://x.com/ankitcode99/status/2108056887363350935
ankitcode99's version of the same point, with a trap: in a 15-turn agent loop about 90% of what you send is text the model already processed, which prefix caching serves at 10% of the normal rate; switch models mid-session and every token recomputes at full price. Session pinning fixes it.
πŸ’‘#18
@brewkeghq
https://x.com/brewkeghq/status/2107860235670958482
brewkeghq, a metered inference gateway, explains why two people on the same plan get a month versus a weekend: every provider gives a monthly pool and a five-hour window, and the variable is the harness. A heavy Claude Desktop coder on Opus 5 burns tokens in hours where someone building the same project in Zed, Pi or a lightweight harness does not, because Claude Desktop ships a large system prompt and tool schema on every turn.
πŸ’‘#19
@_TarunKathuria
https://x.com/_TarunKathuria/status/2107652051497107717
_TarunKathuria's dissent on autoresearch for optimization: not with Fable 5.1 or Astra, at least. Track 3 of modded nanoGPT is not a great optimization benchmark, and even there the autoresearch loop solutions are uninspiring; on well-designed optimization benchmarks they mostly do not work.
πŸ’‘#20
@artrockalter
https://x.com/artrockalter/status/2108309456987795743
artrockalter's one-line critique of selection by metric: the Copernican model was not more accurate than epicycles until Kepler, and so would not have been selected by an autoresearch loop.
πŸ’‘#21
@ant_fitch
https://x.com/ant_fitch/status/2108252933750137331
ant_fitch flags provenance on a widely shared prompt image: it is a harness meta-prompt by another author, adapted from Karpathy's autoresearch project for Claude Code and autonomous loops, not something Karpathy wrote, and it is susceptible to injection attacks.
πŸ’‘#22
@Kyson0531
https://x.com/Kyson0531/status/2108001333986861415
Kyson0531 asks the question the Haiku launch invites: the 75% price drop matters most for the boring jobs, summaries and compaction, which is exactly where numbers drift. In B2B quoting, a compaction that turns MOQ 500 into around 500 is a real bug. Has anyone tested Haiku 5.5 on keeping exact figures through compaction?
πŸ’‘#23
@heyjo_Z
https://x.com/heyjo_Z/status/2108057760676159843
heyjo_Z suggests the same mechanism behind a safety result explains plain task drift: long tool outputs push the original ask out of focus, so re-inserting the request right before the final answer is worth trying in any long agent loop.
πŸ’‘#24
@Policifyai
https://x.com/Policifyai/status/2107913761856360545
Policifyai adds the model split from OpenResearch: Opus 5.5 leads autoresearch usage at 34.9%, GPT-6.1 Sol is a distant second at 25.4%, and open models are only 7.8% of total usage despite their adoption in coding.
πŸ“‘ Eco Products Radar
Eco Products Radar
Claude Code: the harness under most of the loops above, and the thing being optimized in the lead case.
Karpathy's autoresearch: the recipe being adapted to bills, codebases and MAP-Elites outer loops.
OpenResearch: the platform whose usage numbers put Opus 5.5 at 34.9% and DeepSeek at 60.3% of the open share.
Codex: the review leg in the heterogeneous-review paper and the harness blamed for bad GPT runs.
Haiku 5.5: the cheap model everyone is assigning to compaction, QA and lookups.
Pi: the harness used for custom extensions and fine-grained loop control.
← Previous
Super User Daily: 2026-10-09
Next β†’
Ideas Radar: 2026-10-09
← Back to all articles

Comments

Loading...
>_