October 4, 2026loop

Loop Daily: 2026-10-04

The loudest lesson this round came from failures, not wins. An autoresearch agent found a model change that looked great until the dataloader started replaying old text, at which point training loss kept falling while validation stalled, and freezing the suspect memory tables still left 94% of the gap in place. A long essay on recursive self-improvement made the same point from the product side: getting a loop to run is no longer scarce, judging whether it actually improved is, and the fixes that work all keep the final exam out of the loop's reach. Anthropic's applied AI team put it bluntly in a workshop: self-evaluation is a trap, use an adversarial evaluator in its own context window. On the practical end, an agent loop hid a half-finished feature behind a flag and then stopped to ask about blog posts, terms and menus it was not authorized to rewrite, a weekend loop now sorts receipts before Monday, and a Chinese startup says it built a full decision-model family in three days with its own auto-research pipeline. Claude Code mods also pulled the autoresearch loop itself into the editor.
πŸ’‘#1
@SiMin55765
https://x.com/SiMin55765/status/2106422931173879909
SiMin55765 wrote a long essay arguing that in recursive self-improvement the scarce role is the exam writer, not the test-taker. It walks through the evidence: Weco's AIDE run produced seven successively stronger versions over eight days but did not get better at improving, Google's RRSI paper found unconstrained harness search hit 92.8 on its training set and 40.3 on unseen tests against a 39.7 baseline, and four rules, one change per round, a log where dead ideas stay dead, a pre-score review that strips leaked answers, and cost that must be earned, kept transfer gains real with about 30% fewer tokens. The author's own test was letting a model judge interview write-ups: thirty interviews, zero accepted, scores stuck between 0.40 and 0.55 after rewording, until rules took the obvious cases and a human confirmed the middle with one click. The closing advice is to write the success test before building, keep an acceptance set nobody tunes on, and stop when the full cost including human review beats the expected gain.
πŸ’‘#2
@KostyaAI
https://x.com/KostyaAI/status/2106419821415915787
KostyaAI summarized a paper on a failure mode called forking that an autoresearch agent walked straight into. The agent found a model change, an n-gram memory branch, that improved training and validation loss together on the first pass, but once the dataloader began replaying the same corpus, training loss kept falling while validation stalled or rose, with the gap widening at each epoch boundary and rare contexts contributing most of it. Freezing the memory tables did not fix it: in one controlled comparison 94% of the gap remained because the backbone kept fitting the stored representations, and the effect reproduced in a DeepSeek-style architecture. The practical checklist for anyone running automated experiments is to log every dataset pass and replay boundary, group validation by context frequency, compare promising changes with the suspect path removed, and rerun agent-found wins on fresh data and multiple seeds.
πŸ’‘#3
@__MateN
https://x.com/__MateN/status/2105907114677825984
__MateN asked an agent loop to hide a half-finished tax reports feature in the product Tallyroot, and the loop did the in-scope part and stopped at the boundary. It put the feature behind a flag and merged, then parked everything else and asked: three blog posts still mention tax reports, rewrite, unpublish or leave them; the terms have a tax reports section, keep, reword or remove; the roadmap still promises tax reports beyond Hungary; and the settings menu still says Investments and Tax. None of those were part of the task, and the author did not want the agent editing the terms alone, while admitting some of those mentions had been forgotten entirely. One comment on the issue answered all four in two minutes, and the next round picks it up.
πŸ’‘#4
@silenceinwater
https://x.com/silenceinwater/status/2106369875765542956
silenceinwater highlighted a 75-minute AI Engineer workshop where Anthropic applied AI engineers Ash Prabaker and Andrew Wilson called self-evaluation very much a trap. Instead of asking the same agent to build and judge its own work, Anthropic separates planner, generator and evaluator into independent context windows. The evaluator launches the app, tests it with Playwright and returns concrete failures rather than approving its own output. That separation is how they describe keeping agents on track for five to six hours or more.
πŸ’‘#5
@dlimeng192048
https://x.com/dlimeng192048/status/2106351350137455057
dlimeng192048 reported that StartLux built its open-source decision model, StartLux-Decision, in three days using its own Auto Research pipeline. It scored 63.88 on Decision Index 0.2.1 against Jev's 57.91, won 31 of 38 benchmarks, and won 35 of 36 chess games against Jev 1.13. It ships in five sizes from 0.8B to 27B plus quantized local versions, and the 4B model answers three questions in 26 ms on one H200. The company's bet is that high-frequency small decisions become a dedicated compute layer inside local agents, and it describes its pipeline as already at RSI level 3.
πŸ’‘#6
@StartLuxAI
https://x.com/StartLuxAI/status/2106361469126533237
StartLuxAI also released StartLux-V1.0-27B-Preview, a dense model for local agents post-trained on Qwen3.6-27B with what the company calls an auto-research approach. The training focused on verifying intermediate results, adapting to tool feedback and recovering from errors rather than raw knowledge. In an evaluation by CAICT, a research institute under China's industry ministry, using the MCP-Universe method, it scored 39.25% overall micro pass@1, second of six models, ahead of the 284B DeepSeek-V4-Flash and 1.3 points behind the 1.6T DeepSeek-V4-Pro. If the numbers hold, it is a case of an automated research loop producing a small model that competes on agentic tool use.
πŸ’‘#7
@emolyonx
https://x.com/emolyonx/status/2106498202056987104
emolyonx described the economics behind SF Autoresearch, a compute service built for autoresearch runs. You can use hundreds of GPUs for a single run without reserving them in advance and scale up and down across GPUs and CPUs without lock-in. The service is designed to use GPUs more efficiently, so most customers spend several times less for the same result, roughly $100 instead of $500 per 8xH100 node per day. InfiniBand arrives on Monday and B300s later this month.
πŸ’‘#8
@popcoornft
https://x.com/popcoornft/status/2106291426594034169
popcoornft spent a Saturday wiring a small agent loop that watches the downloads folder, renames receipt files by vendor and date, and drops them into the right month folder. Monday used to start with 20 minutes of dragging and renaming. Now the pile is already sorted before the laptop opens. It is the smallest kind of loop, one trigger, one rule set and one destination, and exactly the kind most people never get around to building.
πŸ’‘#9
@ProductLogAI
https://x.com/ProductLogAI/status/2106479281261957179
ProductLogAI spent four hours debugging an agent loop before finding that the frontier model was not failing any logic test at all. It was stuck in a passive retry state because an edge-case API response returned null instead of an empty array. The loop treated the null as something to wait on rather than as an answer. It is a reminder that many loop failures live in the tool contract, not the model.
πŸ’‘#10
@AgentonomyXYZ
https://x.com/AgentonomyXYZ/status/2105827234485551392
AgentonomyXYZ spelled out one design rule for payment loops: when a payment times out, the agent should not simply pay again. The next step is to check the original request, because payment and delivery are tracked separately. An efficient loop recovers from uncertainty without repeating the charge. It is a small rule that addresses one of the most expensive failure modes for agents near money.
πŸ’‘#11
@Mersad_Abbasi
https://x.com/Mersad_Abbasi/status/2106213749636145403
Mersad_Abbasi offered a two-question test before starting any autoresearch experiment. Does the agent have a large solution space to search, and have you built a tiny objective interface to evaluate it? The author argues that every successful test-time scaling experiment has both. Without the second, a loop has nowhere to point its effort.
πŸ’‘#12
@EGafni
https://x.com/EGafni/status/2106239570354553267
EGafni shared two favorite patterns that combine autoresearch with decision models. In autoresearch for feature extraction, an LLM generates questions for Jev whose answer probabilities are fed into a classic ML model such as logistic regression to predict a label, so the loop searches over questions rather than weights. In hierarchical classification, Jev traverses a taxonomy with beam search, which the author says it is especially good at. Both turn a fast yes-or-no model into a component the loop can optimize around.
πŸ’‘#13
@haydonryan
https://x.com/haydonryan/status/2106440027404218420
haydonryan proposed pointing an autoresearch loop at rustc. On an EPYC 7443 the longest step is linking, and a cursory scan suggests the compiler is already efficient, so the next idea is a rewrite-in-assembly experiment that uses the Rust test suite and asm blocks to see how fast it can go. The author notes the real cost is running extensive test suites on complex apps: you can get away with fewer tests, but that risks bugs. The preference is more checks over dealing with failures, which is exactly the trade an autoresearch loop has to encode.
πŸ’‘#14
@stretchcloud
https://x.com/stretchcloud/status/2106020376186962279
stretchcloud read a viral 18-item checklist for agent builders, from durable state machines and token budgets to trajectory grading in CI and a 50-cent-per-query kill switch, as proof that agent reliability is now a market. Temporal just raised $550 million at a $12.55 billion valuation because agent workflows need crash recovery a while loop cannot give, and observability split into its own stack: Langfuse self-hostable from free, LangSmith at $39 a seat, Braintrust at $249 a month, Arize Phoenix open source with a $50 hosted tier, Helicone at $20. None of them sell a smarter model; they sell the replay log, audit trail and circuit breaker. The author's observation is that everyone ships an agent that reasons well, but far fewer build the switch that stops a loop from spending real money before anyone notices.
πŸ’‘#15
@eveliqTrace
https://x.com/eveliqTrace/status/2106052493180293221
eveliqTrace argued that a faster agentic loop is not the same as a governed one, infrastructure capability is not task-specific authority, and an executed action is not a verified effect. As agents call tools, run code and change external systems, the chain has to extend past compute to identity, authority, action, observed state change, verified effect and revalidation. Infrastructure can accelerate the loop, but assurance has to show each material transition was authorized, attributable and produced the intended result. The post frames that as the point where agent infrastructure and evidence infrastructure need to meet.
πŸ’‘#16
@privacymage
https://x.com/privacymage/status/2106431602930671881
privacymage is spending the weekend on an autoresearch competition run by Yukon Research with the Ethereum Foundation, aimed at Ethereum's post-quantum future. The entry uses agentprivacy's dual-agent harness, sig_mage, and the author frames it as sharing a spellbook with other agents along the way. It is a sign that open autoresearch competitions are becoming a format, with public problems, shared harnesses and agents contributing to the same target.
πŸ’‘#17
@zeroknowledgefm
https://x.com/zeroknowledgefm/status/2106397396385099958
zeroknowledgefm released an episode where Yukon Research's Soubhik Deb describes testing whether multiplayer auto research, the idea behind an earlier one-off result, would work on a second problem. The question is whether a crowd of agents and people attacking one public problem was a fluke or a repeatable method. The episode walks through what happened on the second attempt. It is the methodological follow-up the format needed.
πŸ’‘#18
@hevmind
https://x.com/hevmind/status/2106141280313242046
hevmind described the architecture of hev ask, which skips the vector database entirely. Claude builds a digest of a documentation site offline, and at query time a constrained agent loop gets at most four tool calls to answer straight from the site's own pages. Bounding the loop at query time keeps cost and latency predictable while still letting it look things up. It is a neat example of putting the expensive thinking offline and keeping the live loop small.
πŸ’‘#19
@tichmangono
https://x.com/tichmangono/status/2106194144653639928
tichmangono is building a black-box software factory and asked who else is. The stack is herdr and Pi, following proper SDLC steps in a loop, and borrowing graph engineering principles plus pieces of Karpathy's autoresearch. The point is that the loop runs the whole lifecycle, not only code generation. The author offered to share the repo with anyone interested.
πŸ’‘#20
@VibeCoderOfek
https://x.com/VibeCoderOfek/status/2106195908081676480
VibeCoderOfek pointed out that middleware inside the agent loop is the part most teams skip, bolting a plugin on after the run instead. The ask is for a mod that sits before the tool call and denies the write if the check is red. Having Claude write that mod is fine, the author says, but letting the mod ship itself is not. It is a precise line between using the agent to build its own guardrails and letting it approve them.
πŸ’‘#21
@The_Tradesman1
https://x.com/The_Tradesman1/status/2106303672770818155
The_Tradesman1 described a community Claude Code mod that runs the autoresearch loop inside the editor: try an idea, benchmark it, keep the win, roll back the regression, repeat. It is an almost one-to-one port of pi-autoresearch built on the new mods API, and sessions use the same .auto/log.jsonl format, so logs can be swapped between the two tools. The demo copies an example sort benchmark into a git repo and runs /autoresearch make sort.js faster, and it needs Claude Code 2.1.285 or newer with function hooks enabled in settings. The question the post leaves open is the right one: would you let a loop decide which of your changes survive?
πŸ’‘#22
@ashtewari
https://x.com/ashtewari/status/2105848131795488947
ashtewari took the automated agentic loop that Karpathy's autoresearch applies to model-training research and applied it to fine-tuning instead. The experiment and its write-up are linked in the post. Moving the loop from pretraining speedruns to fine-tuning puts it on the kind of task more teams actually run. It adds one more domain to the list of places the propose, run and keep pattern is being tried.
πŸ’‘#23
@kylemissionai
https://x.com/kylemissionai/status/2106445461749997673
kylemissionai argued that everyone is building the same agent loop because the loop is table stakes and will converge the way databases did. The real differentiation moves up a layer, to who the user trusts to act on their behalf when they are not watching. In the same thread sachintwtss put it as the moat being the boring permissions and failure recovery that nobody demos. Both point at the same gap Loop readers keep seeing: the loop is commoditizing, the trust layer around it is not.
πŸ“‘ Eco Products Radar
Eco Products Radar

Jev: the decision model showing up inside loops as a feature generator, a router and the benchmark that new decision models measure against.
Claude Code Mods: the hook layer that now carries loop logic, from an in-editor autoresearch port to guards that deny writes before a tool call.
Pi and pi-autoresearch: the minimal harness and the loop recipe that the Claude Code mod ports one-to-one.
StartLux: two releases in one window, both credited to its own auto-research pipeline.
Langfuse and Temporal: the observability and durable-execution layers named whenever a loop has to survive real traffic.
Andrew Ng's agentic AI course: shared by four separate accounts, mostly for its section on self-improving agent loops.
← Previous
Super User Daily: 2026-10-04
Next β†’
Ideas Radar: 2026-10-04
← Back to all articles

Comments

Loading...
>_