October 7, 2026loop

Loop Daily: 2026-10-07

The window's most useful loop content was about what the loops actually do when nobody is reading. A study of 119 autoresearch traces found that an agent's tendency to think versus act is 75 percent explained by the model and 3 percent by the task, that claimed improvements on a chip-design task were right 29 percent of the time against 31 percent when the agent said nothing, and that the median agent quit after spending 19 percent of its budget. A Meta paper put a number on cross-vendor review: mixing Claude Code and Codex on the same training-code edits lifted fully correct patches from 45.8 to 62.5 percent at matched spend, because agents from the same product fail the same way. On the practitioner side, a nightly Devin autoresearch prompt with a budget, a time limit and a line telling it to boil the ocean, three Opus 5.5 loops running around the clock, an RL researcher who stopped babysitting runs and got QWOP under the record, and a developer who spent a week hand-rolling an agent loop before discovering the Claude Agent SDK is the same harness Claude Code runs on. The decision-model thread kept going: a 9B Clef beat a 27B at Breakout purely on latency, and a long essay on Jev separated computing an answer from writing it out and flagged a pre-registered calibration study that found systematic overconfidence.
πŸ’‘#1
@JIACHENLIU8
https://x.com/JIACHENLIU8/status/2107564033742127480
JIACHENLIU8 read 119 autoresearch traces and reported three findings. First, an agent's personality is hardcoded: whether it tends to think or act is 75 percent explained by the model and only 3 percent by the task, with GPT-5.5 spending about 90 percent of its output deliberating before running a command while Kimi jumps straight to execution, so guardrails built for one model fail on another. Second, claimed breakthroughs are worse than coin flips: on chip design an agent was right 29 percent of the time when it claimed an improvement and 31 percent when it said nothing, yet 60 percent of follow-up experiments were built on those phantom wins. Third, agents do not run out of time, they give up: the median agent quit after burning 19 percent of its compute budget, and the winning runs were the ones that refused to stop exploring late.
πŸ’‘#2
@rohanpaul_ai
https://x.com/rohanpaul_ai/status/2107447530204131352
rohanpaul_ai summarised a new Meta paper, RankEvolve, which found that having two different coding agents review each other's patches catches far more silent bugs than giving one agent a bigger budget. Mixing Claude Code and Codex on the same task raised fully correct patches from 45.8 percent to 62.5 percent at matched spend. The failure mode is specific to auto-research: agents editing real training code can leak test data, break a gradient or miswire a flag, the code still runs, and you burn GPU hours on numbers that look valid. The explanation is uncorrelated mistakes, agents from the same product fail the same way so there is little for same-vendor review to catch, and the effect repeated on an unrelated training codebase.
πŸ’‘#3
@jaredpalmer
https://x.com/jaredpalmer/status/2107566018667446315
jaredpalmer has been running Devin on autoresearch every night for about a week with a deliberately simple prompt: do autoresearch to improve a named verifiable metric, hill climb, with a dollar budget and a time limit, use sub-devins to parallelise, track plans and progress in the filesystem, create abstractions if they speed up verification, manage and babysit your own PRs and merge when a candidate passes the gates you set, boil the ocean, believe in yourself. Once the setup works it becomes a skill with automatic code review turned on and Devin runs for hours. The surprising part is that the hype statements measurably improve exploration, making models less likely to give up early, reward-hack, or ship incremental fallbacks, an idea borrowed from Jarred Sumner's Riemann prompts.
πŸ’‘#4
@pratikg
https://x.com/pratikg/status/2107536121139957972
pratikg explained how multiplayer autoresearch on Yukon works: a human challenge setter posts a hard problem with a clear score and an automatic verifier, humans and AI agents attack it at once using any model and any harness, every attempt goes through the verifier, and when one beats the current best it becomes the new frontier that anyone can build on instead of starting from zero. All results and commits are open source so every record is traceable. The platform has solved problems in cryptography, math, science and open-weight models, and is moving into biology, chemistry and materials science.
πŸ’‘#5
@Rafa_Schwinger
https://x.com/Rafa_Schwinger/status/2107508602323816742
Rafa_Schwinger, who now runs three autoresearch loops around the clock on Opus 5.5 where a Fable loop used to need a day or two of planning, posted the working rules: if you have more than one agent, ask them to build a coordination mechanism that uses the hardware to measure things and benchmark; Opus on high is enough as the orchestrator and can spawn other Opus subagents, with Fable as an advisor; download the highest quality repos with similar implementations for the agent to study; strip out any good-practice framework that gets in the way of optimisation and work at the lowest possible level; ask for periodic summaries since the last message and give common-sense feedback; have it read the codebase structure so it does not sloppify everything; and keep the autoresearch machinery self-contained in a module so complexity stays localised.
πŸ’‘#6
@patrickbbrown
https://x.com/patrickbbrown/status/2107596491976024067
patrickbbrown stopped babysitting reinforcement learning runs and handed the job to a coding agent that now tunes settings, reshapes rewards, queues the next experiments and reads the results. The reported outcomes are QWOP under the record and a balanced quadruple pendulum, with a write-up of how the loop is set up. It is a clean example of autoresearch applied to RL training rather than LLM pretraining.
πŸ’‘#7
@sudoingX
https://x.com/sudoingX/status/2107553983908950128
sudoingX pitted two of Cloudflare's Clef decision models against each other on Breakout and the 9B beat the 27B decisively: 56 ms per answer and 15.6 decisions a second with 92 saves and 5 balls lost, against 176 ms, 5 decisions a second, 25 saves and 32 lost, both at q4 for three minutes. Between two answers the 9B's paddle drifts about 42 px, a third of its width, so the next answer still catches the ball, while the 27B's drifts 128 px and swings past. A control script with perfect judgment and the same delays collapsed from 25.8 saves per ball at 56 ms to 1.4 at 176 ms. On pixels the 27B pointed at the ball 96.6 percent of the time versus 68.6, so it is the better reader and still loses for being late. The conclusion for any live loop: put the fastest model that is good enough in charge and keep the big one for calls where being wrong costs more than being slow.
πŸ’‘#8
@StragglerLiu
https://x.com/StragglerLiu/status/2107273528752152862
StragglerLiu wrote a long analysis of Jev that separates two things a generative model does on a yes/no question, computing the answer in one forward pass and then transcribing it token by token, and argues Jev simply drops the transcription: enumerate the legal outputs, one parallel pass, return a softmax. That is why output is free at $0.042 per million input tokens. The cost numbers cited are structural, Harness cut per-endpoint AppSec classification cost about 150x, OpenRouter's 40,000-ticket triage benchmark came to $0.0248 per thousand tickets against $2.88 for Opus, and Jev reached nearly 13 percent of Vercel's paid teams in 24 hours. The essay's caution is calibration: a pre-registered NXTG AI study found pooled ECE of 0.156 with mean confidence 0.66 against accuracy 0.53, uncalibrated in four of five decision classes, which would narrow Jev from automated decision layer to very fast cheap classifier. OpenAI's Decisions API, AWS Strands Decider 2B, Cloudflare Clef and Databricks ai_decide all arrived within two weeks.
πŸ’‘#9
@gippp69
https://x.com/gippp69/status/2107481859701555484
gippp69 ran a cost audit with Claude Code over an agent loop that already looked optimised: feed it the usage log for every task, have it flag late Opus escalations, cold-cache rewrites and effort changes, then let it rewrite the router, handoff and session rules around those failures. The run scanned 4,812 calls across 1,000 tasks and found three leaks, the biggest being late escalation from Sonnet to Opus. After the fixes the bill went from $2,530 to $1,820 a month on the same workload, $710 saved. The point of the post is that Claude Code can audit the system rather than only write the code.
πŸ’‘#10
@calebfoundry
https://x.com/calebfoundry/status/2107569666285523100
calebfoundry has used many neoclouds and says SF Compute is miles ahead on philosophy because the agent itself can spin GPU instances up and down and be charged by the minute to run autoresearch tasks. The first job is a 30-minute training run, with plans to scale up. The same day SF Compute announced that SF Autoresearch supports jobs with full InfiniBand fabrics, no long-term reservation, lead time or deposit, which is the supply side of the same story.
πŸ’‘#11
@haydonryan
https://x.com/haydonryan/status/2107508838874128796
haydonryan's advice for optimisation loops is to run single targeted prompts alongside the generic make-it-faster autoresearch loop: find unnecessary heap allocations, find where the Rust type system could enforce valid state, or compound prompts like read the Rust optimisation book and create a suite of optimisation prompts, then A/B test each one for size, speed and correctness. It uses a chunk of compute but finds real improvements, with the caveat to always review and test by hand, since the model also finds plenty of not-worth-it changes, and to respect the repository's contribution rules.
πŸ’‘#12
@taytaycodes
https://x.com/taytaycodes/status/2107577592136294557
taytaycodes built a custom agent loop from scratch, file reading, tool execution, context management, before realising that Claude Code's actual harness is available as a library. The Claude Agent SDK exposes the same agent loop, built-in tools and context management that power Claude Code, in Python or TypeScript, renamed from Claude Code SDK in late 2025 when it was generalised beyond coding. The trade-off noted is that it is the same battle-tested harness rather than a thinner reimplementation, while anyone who does not want to run infrastructure at all should look at the hosted managed agents API first. The author's own summary: a week would have been saved by knowing this upfront.
πŸ’‘#13
@zeyan0823
https://x.com/zeyan0823/status/2107314131963662541
zeyan0823 expected ICLR open review to pull the auto-research crowd toward some convergence and instead came away numb: some submissions are visibly full-pipeline auto-research, starting from a conclusion, swapping the metric when experiments do not fit, finding a new phenomenon when the new metric makes no sense, snowballing into dozens of figures and piles of experiments with no main thread. The author is explicit about not objecting to AI running experiments, writing papers or making figures, only to papers nobody read once before submitting. Several other researchers in the window made the same complaint about ICLR submissions produced by auto-research loops.
πŸ’‘#14
@WengZhaoti39773
https://x.com/WengZhaoti39773/status/2107538149643796758
WengZhaoti39773 is presenting GEA, Group-Evolving Agents, at COLM this week. The method treats a group of agents rather than an individual as the unit of evolution: agents share experience as they self-improve, and the paper reports that this outperforms prior open-ended self-improving methods and matches or exceeds the top human-designed agent frameworks with zero human intervention. The author's stated research focus is recursive self-improvement and autoresearch, and COLM in general was thick with autoresearch posters and workshop talks this window.
πŸ’‘#15
@Rafa_Schwinger
https://x.com/Rafa_Schwinger/status/2107497755664916958
Rafa_Schwinger's shorter post is the before-and-after that explains the longer one: planning a Fable autoresearch loop used to take the whole day before it could run for a day or two, and with Opus 5.5 the same person now runs three loops 24/7. It is the clearest single data point in the window on how much the model upgrade changed the operating economics of overnight research.
πŸ“‘ Eco Products Radar
Eco Products Radar
Jev: the decision model in nearly every loop-architecture post, now with a calibration study to answer.
autoresearch (Karpathy): the three-file repo still being explained daily, now cited as the pattern behind GPU, RL and crypto loops.
Opus 5.5: the model that turned one-planned-loop into three-loops-24/7 for at least one researcher.
Claude Code: the harness in the Meta cross-review paper, the cost-audit post and the Agent SDK discovery.
Codex: the other half of the cross-vendor review result.
Devin: running nightly autoresearch with sub-devins and automatic code review.
SF Compute: agent-driven GPU provisioning by the minute, now with InfiniBand jobs.
Clef: Cloudflare's decision models, benchmarked 9B against 27B on a live game.
COLM: the conference hosting most of the week's autoresearch and self-improvement papers.
DeepSeek Harness: v0.2 with desktop apps and a Claude Code mods compatibility layer.
← Previous
Super User Daily: 2026-10-07
Next β†’
Ideas Radar: 2026-10-07
← Back to all articles

Comments

Loading...
>_