October 2, 2026loop

Loop Daily: 2026-10-02

The autoresearch story this window has a price tag on it. Sam Hogan published the bill for a Kimi K3 inference-optimization loop: 120 B300s, six hours, a 40% efficiency gain, about $9,200 all in, paid back in under 12 hours. The research side caught up the same day with a loop that builds its own benchmarks and finds that humans matter most exactly when the loop stalls, plus a tool that lines Codex agents up against Kaggle Grandmasters and feeds the humans' habits back in. The counterweight is just as loud: the person who set the latest NanoGPT record says autoresearch played almost no role, and a well-known researcher points out that a pile of auto-research runs missed an optimization that looks obvious in hindsight. Underneath it all, the practical loop engineering got sharper: where the timestamp sits in your prompt decides whether 50 agent calls cost $6.25 or $0.37, and the hardware signer, not the model, is what stopped an agent from paying ten times too much.
πŸ’‘#1
@samhogan
https://x.com/samhogan/status/2105142764258431419
samhogan rolled out a Kimi K3 inference-optimization autoresearch loop that ran on 120 B300s and found a 40% efficiency gain in under six hours. The bill was itemized: compute at 120 GPUs times 6 hours times $5.80 an hour is $4,176, engineering plus token cost about $5,000, total $9,176. At current serving scale it pays for itself in under 12 hours. This is the cleanest public ROI number for autoresearch on production infrastructure so far, and it explains why inference teams are adopting the loop first.
πŸ’‘#2
@jaseweston
https://x.com/jaseweston/status/2105305463784935791
jaseweston introduced AutoBenchmark, which creates benchmarks automatically and benchmarks the creation process itself. The key finding is that human-agent collaboration beats agents alone, with fine-grained human feedback at the ideation stage mattering most, and that benchmark creation works best with feedback from two sources, the solvers and external verifiers both human and AI. They also show it is possible to build autoresearch benchmarks for AI research with this recipe, a full recursive improvement loop. The last post in the thread adds the most useful detail: when the autoresearch loop stalls, human direction at that point is what helps, tested on a new Rebuttal Bench that judges whether a paper rebuttal actually resolves reviewer concerns.
πŸ’‘#3
@sunweiwei12
https://x.com/sunweiwei12/status/2105405100172722496
sunweiwei12 introduced TraceML, a tool for trajectory-level analysis of research agents and systematic agent-versus-human comparison, to be presented at NeurIPS 2026. Comparing Codex agents against Kaggle Grandmasters, they found humans draw on a broader set of research skills while agents concentrate on a narrow subset. The interesting step is the reverse transfer: distilling the human research skills that TraceML uncovered back into Codex agents substantially improved final outcomes. It turns the vague claim that agents research differently into something you can measure and patch.
πŸ’‘#4
@prathamgrv
https://x.com/prathamgrv/status/2105234888555385126
prathamgrv solo-authored a research paper on an autoresearch problem they had been curious about for months, and it was accepted at the NeurIPS AutoMLR workshop. They ran a ridiculous number of experiments, burned about $1,000 of H100 time, and published under their own company's affiliation. A single person producing a workshop paper on autoresearch for about a thousand dollars of compute is itself a data point about how cheap this kind of research has become.
πŸ’‘#5
@chuanyang_jin
https://x.com/chuanyang_jin/status/2105313263701598333
chuanyang_jin, a co-author on the automatic benchmark study, framed automatic benchmark creation as one of the most open-ended, high-value directions for autoresearch agents. The point behind it is that a loop can only optimize what it can score, so the scarce input is not more agents but more and better targets. Benchmarks written by loops, checked by solvers and outside verifiers, are how autoresearch extends into domains that do not come with a metric.
πŸ’‘#6
@nasqret
https://x.com/nasqret/status/2105079853506769407
nasqret has spent the past few weeks trying a new kind of auto-research loop for mathematics. The goal is not to solve one specific problem but to bootstrap the shape of the best possible result: what is actually provable, what is false, and what is still dark. In a follow-up they stress that each project needs a different setup, that using the same rigid configuration everywhere is very ineffective, and that doing it right is a skill in itself requiring real expertise in the field. It is a useful corrective to the idea that the loop replaces the mathematician.
πŸ’‘#7
@dzhng
https://x.com/dzhng/status/2104405629993676961
dzhng released jevgrep v0.4 after another 24 hours of autoresearch loops, reaching intelligence parity with Codex's and Claude's own subagents at 30% less cost. The previous version saved 40%, so this release deliberately traded some savings for more intelligence. Letting a loop tune a developer tool against a parity target, then publishing the tradeoff honestly, is a practical template for shipping loop-optimized software.
πŸ’‘#8
@ethereumfndn
https://x.com/ethereumfndn/status/2105314932283814382
ethereumfndn launched an open autoresearch challenge: design stateless hash-based signatures for post-quantum Ethereum, prove them secure in Lean, and beat the record on signature size times verification cycles. Participants point their agents at the challenge site, built with Eigen Labs on Yukon Research. A foundation publishing a formal-proof-gated leaderboard for agents is a sign that cryptography is becoming a natural autoresearch domain, because both the score and the correctness check are machine-verifiable.
πŸ’‘#9
@JackieLeeETH
https://x.com/JackieLeeETH/status/2105131787748032804
JackieLeeETH added a playback mode showing how a multiplayer autoresearch effort broke records over the past few months while still optimizing the SECP256K1 point-addition circuit. The interactive chart lets you replay the record progression. Watching many independent agents and people push one circuit down step by step makes the shared-leaderboard model of autoresearch tangible.
πŸ’‘#10
@adi_baradwaj
https://x.com/adi_baradwaj/status/2105066306798260558
adi_baradwaj passed on a notable counterpoint: Deven, who set the recent NanoGPT speedrun record, says autoresearch played almost no role in the work. All ideation was done by a human and implementation was done by Claude Code. In other words the coding agent mattered and the research loop did not, at least for a record that rewards novel ideas. It is worth keeping next to the big ROI numbers, because the two are not in conflict: loops shine at tuning, humans still own the leaps.
πŸ’‘#11
@sytelus
https://x.com/sytelus/status/2104815472705409470
sytelus praised a set of techniques that skip compute for token IDs absent from a typical batch, which also improve inference performance, and then made the sharper point: tons of auto-research efforts had not found these optimizations, which feel very obvious in retrospect. Their conclusion is that humans still have an enormous creative edge. Loops search well inside a frame; changing the frame is still mostly a human move.
πŸ’‘#12
@int21_ai
https://x.com/int21_ai/status/2104602311368610093
int21_ai reported that two engineers directed the generation of 20 inference engines in two weeks across seven model categories: autoregressive and diffusion language, speech recognition, speech generation, OCR, image and video generation. Benchmarked on a 15-model subset against SGLang and vLLM, INT21 led on decode rate for all six text models tested, with MiMo at 1,308 tokens per second against 540 for tuned SGLang and 1,011 for vLLM, and complete audio generation came out 1.73 to 2.30 times faster than SGLang-Omni on the same GPU allocation. They credit self-improving agent swarms for building the systems that run AI faster, and argue that in the agentic era inference latency becomes business latency.
πŸ’‘#13
@bendechrai
https://x.com/bendechrai/status/2105078234513551810
bendechrai traced 14 months of building software factories: first a humble Ralph loop, then IssueOps running entirely in GitHub Issues and Actions until API billing hurt, then Anneal on a local Claude subscription wrapping Ralph loops with outer loops and tickets, and finally Holodeck, where projects are built inside the factory and the first project it built was itself. The lessons are concrete: deterministic gates outside the agent loop are the real quality mechanism, the same agent never writes and grades its own work, fresh contexts always with tickets logging what failed, every editing agent gets its own worktree in throwaway containers with no Docker socket, and agents get references to secrets, never values. The spec is now the hard part, and tools are disposable because every model release makes some of them redundant.
πŸ’‘#14
@0xbobaaa
https://x.com/0xbobaaa/status/2105682175035138331
0xbobaaa showed that the same 50-call agent loop on GPT-6.1 Sol costs $6.25 or $0.37 depending only on where one timestamp sits. Sol reads cached input at $0.10 per million tokens, 95% off, and keeps it warm for 30 minutes, but the cache matches from the first token, so a timestamp at the top makes every call a fresh write at 1.25 times. For a 50,000-token prompt resent 50 times: timestamp first is $6.25, no cache at all is $5.00, timestamp last is one write plus 49 cheap reads at $0.37, and the same loop on Astra without cache is $25. The check is simple: look at cached_tokens on your second call, and if it is zero, something near the top moves.
πŸ’‘#15
@sairamkotha
https://x.com/sairamkotha/status/2105584235780493559
sairamkotha made the same point from the other side: every time you evict early turns in an agent loop to save context, you invalidate your prefix cache. A 90% prompt cache discount drops to zero, p95 time to first token quadruples and token spend surges. The architectural fix is append-only logs with isolated scratchpads. A related post argues context inflation is a systems problem where dynamic cost forecasting beats static step caps.
πŸ’‘#16
@themahis
https://x.com/themahis/status/2104629088623165725
themahis told an agent to buy one pair of headphones; it took the cart to checkout as ten and sent a 0.10 ETH signing request to a Ledger instead of 0.01. On the Speculos emulator they opened the review screen and rejected it, the app returned error 0x6985, no signature, nothing left the device and nothing hit the network. That is the whole point of a signer in the agent loop: the agent can propose the cart but cannot spend it. Hardware approval outside the model is the most convincing guardrail demo of the week.
πŸ’‘#17
@deeepakbagada
https://x.com/deeepakbagada/status/2104817449744949676
deeepakbagada described a Rust MCP schema proxy that they say blocked all tool-parameter hallucination crashes across 92,000 daily tool invocations. It validates payloads against JSON Schema in 1.8ms before execution, deterministically coerces stringified integers and ISO dates, strips hallucinated arguments before calling tools, and isolates tool crashes in a process sandbox so the parent agent loop survives. The result claimed is 40-step workflows with zero runtime type exceptions. A typed boundary between the model and the tools is cheap insurance for any long loop.
πŸ’‘#18
@kumarumt
https://x.com/kumarumt/status/2104370512294015269
kumarumt posted the permission fence they set up before any client job on the Claude Agent SDK: start with allowedTools empty and add one tool at a time with a reason, keep money, send, delete and push on human approval forever, name MCP servers explicitly, use hooks on tool start and end for audit, give subagents narrower tools than the parent, treat forking a session without a new permission sheet as a bug, and rehearse denial by asking the agent to use a banned tool. The closing rule is the strongest: if you cannot print who approved each tool, it is not production. The model is not the product, the permission table is.
πŸ’‘#19
@jatingargiitk
https://x.com/jatingargiitk/status/2104929338097549624
jatingargiitk shared a lab-automation case where Opus 5.5 worked out how to clear a 3D printer bed after 16 attempts. Now the printer runs experiments continuously with no human needed. Physical-world loops where the agent must fix its own setup before the real experiment can repeat are where autonomy pays off most, because the bottleneck was always a human walking over to reset things.
πŸ’‘#20
@RMladek
https://x.com/RMladek/status/2105057398566367666
RMladek published how the personal agent Instinct works under the hood: the agent loop runs outside the sandbox on the backend server, the sandbox only executes commands and controls the file system, and a git memory folder is backed up to S3. It ships around 30 narrow skills, from flight booking and food ordering to document signing, Mac computer use and proactive value creation. A follow-up teardown by code_rams summarized the lesson as splitting brain, hands and memory so that when the machine dies, the memory does not.
πŸ’‘#21
@dani_avila7
https://x.com/dani_avila7/status/2105276161525719071
dani_avila7 got Meta's Muse to describe its own architecture with a couple of questions: an agent loop of understand, plan, act, verify and reply, a persistent Linux VM with terminal and internet, a live Chromium browser with persistent logins, connected services from Gmail to health data, background subagents, scheduled jobs and event hooks, memory for preferences, projects, people and commitments, and safeguards around publishing, sending, buying and credentials. They suspect something close to this is in the system prompt and plan to ask periodically to see what changes. It is a free reference architecture for anyone building an always-on agent.
πŸ’‘#22
@hwchase17
https://x.com/hwchase17/status/2104940633018577400
hwchase17 compared a company OS, an org-level harness, with a personal agent. The differences: many principals and one actor, so it must be multiplayer and handle auth and memory correctly even with several users in one thread; governance, observability, auditability and an admin control plane matter far more; and interaction is more event-driven and async. The similarities: it should write and execute code, browser use is likely very important unlike in a coding harness, skills and MCP are the standards, and the core agent loop is shared. It is a useful checklist for anyone deciding whether to adapt a personal agent for a team.
πŸ’‘#23
@evisdrenova
https://x.com/evisdrenova/status/2104609018694222311
evisdrenova asked why everyone obsesses over sandbox startup time when it is the least impactful metric in an agentic loop. Their numbers: sandbox startup is 10 to 50 milliseconds, one tool call is 200 to 500 milliseconds, and one LLM round trip is two seconds or more. Optimize the thing that dominates the loop. The same day CoreWeave announced a CPU platform claiming 3x faster sandbox startup, which makes the question pointed.
πŸ’‘#24
@raahulll_raj
https://x.com/raahulll_raj/status/2105618236326981904
raahulll_raj pointed out a failure mode in trading agents: the agent reads price 100, spread 0.2% and healthy liquidity, reasons for two seconds, calls another model or tool, builds the transaction and sends it, and by then the price is 103 and the market has changed. The decision was right for a state that no longer exists. The fix they propose is an expiry on observations, so the loop becomes observe, reason, re-check state, execute. It generalizes beyond trading to any agent acting on a world that keeps moving.
πŸ’‘#25
@naveenpandey27
https://x.com/naveenpandey27/status/2104923102417469813
naveenpandey27 highlighted a small update in Google's Antigravity CLI 1.2.13: it now respects server-provided retry delays instead of retrying on a fixed timer, and stops retrying when the provider signals a daily or billing quota limit. Their point is that an agent loop should not treat every failure as retry, retry, retry, but distinguish transient failure to retry, rate limit to back off, quota exhausted to stop, and persistent failure to escalate. Good agent systems are not just smarter models, they are better runtimes.
πŸ’‘#26
@HOPPYEMPIRE
https://x.com/HOPPYEMPIRE/status/2105217893860270304
HOPPYEMPIRE builds every agent loop around one question: do we have new evidence for another action? The loop is plan, act, inspect, decide, and that stop condition does more work than another paragraph of prompting. It is a compact answer to the most common loop failure, agents that keep going because nothing told them to stop.
πŸ’‘#27
@EgeMustafaCelik
https://x.com/EgeMustafaCelik/status/2104462443833401444
EgeMustafaCelik argued that the gap between a long agent loop and a hive mind keeps getting flattened in press releases: a number like 10,000 agents over 88 hours can still be one controller with a retry shell. They treat swarm claims as orchestration claims until there is evidence of shared state, failure isolation and a stop condition that is not just more tokens. It is a good filter to apply to the self-improving-swarm announcements now arriving weekly.
πŸ’‘#28
@SqdiqX
https://x.com/SqdiqX/status/2104940456459063533
SqdiqX noted what made the Raven announcement different: the team used recursive self-improvement to ship the thing announcing it, and the orchestration layer is its own instance that can be rewritten too. Most self-improving agent projects stop at prompts or model swaps because touching orchestration means a bug breaks the thing coordinating every task. Letting a harness rewrite its own coordination logic is either well engineered or about to produce some interesting failure threads, and either way it is more informative than another benchmark.
πŸ’‘#29
@pzakin
https://x.com/pzakin/status/2105034454167425228
pzakin listed the ideas they keep coming back to: sims plus sensors as core infrastructure for autonomous decision-making, explorer agents that do autoresearch in non-verifiable domains, and taste replicators or personal models that let agents apply subjective judgment at scale, such as evaluating designs. The second one is the frontier: autoresearch works where there is a score, and the open question is what replaces the score everywhere else.
πŸ’‘#30
@sonicdr1p
https://x.com/sonicdr1p/status/2105339586901586375
sonicdr1p summarized a creator who tested hundreds of Claude skills and kept nine, with measured numbers that came in well below what the repos promise. Caveman claims 65% fewer tokens and measured about 35%, Ponytail claims up to 94% less code and measured around 50%. Karpathy's autoresearch, built for training models, was pointed at a gym plan, and the method of running 100 tiny tests and keeping what works caught that their arms only got 7 sets a week. A non-coding autoresearch run on a workout plan is a small but telling sign of where the method is spreading.
πŸ’‘#31
@cybersnopy
https://x.com/cybersnopy/status/2105415516517003266
cybersnopy reported CoreWeave going first on NVIDIA's Vera CPU, which it calls the first processor designed for AI agents: 128 Vera CPUs and 11,264 cores per rack with BlueField-4 DPUs, enough for more than 11,000 concurrent agent environments. The use case is the CPU half of the agentic loop, sandboxes, RL environments, tool calls and data pipelines around every model step, and that demand is bursty, thousands of environments for an hour and then almost none. CoreWeave claims 3x faster sandbox startup on Vera versus x86. Infrastructure vendors designing around the loop rather than the model is the structural shift here.
πŸ’‘#32
@StragglerLiu
https://x.com/StragglerLiu/status/2105136401247539424
StragglerLiu wrote a long analysis arguing the agent runtime, not tokens, is now the distribution layer. OpenAI's Agents API hosts the Codex harness, compaction, tool scheduling and orchestration but lets execution run in third-party or your own sandboxes with no platform fee, while Anthropic's Claude Managed Agents hosts loop, tools and sandbox together at standard token pricing plus $0.08 per session hour. Citing Goldman Sachs, the same programming task uses 3,390 tokens as a single Q&A and about 4.17 million in an agentic workflow, so whoever holds the state collects the multiplier. The metrics to watch are retention under each model and whether anyone migrates a production agent between runtimes.
πŸ’‘#33
@AlemTuzlak
https://x.com/AlemTuzlak/status/2105583984168050890
AlemTuzlak, who works on TanStack AI, says the project has been underselling itself: its chat method is a fully fledged agent-loop utility that can be used to build a custom harness or anything in between, with all the other functionality on top. Agent-loop primitives moving into mainstream framework libraries means fewer teams will hand-roll the loop, which in turn makes the runtime choices above more about state and governance than plumbing.
πŸ’‘#34
@rasbt
https://x.com/rasbt/status/2104996058040283385
rasbt suggested seeing agent tools on a spectrum: doing everything yourself, then a Codex-like harness for manual automation, then a Dots-like always-on agent loop for full automation. They prefer separating tasks and projects into threads, and putting everything from website maintenance to expense management into one agent chat horrifies them. It is a sensible counterweight to the one-agent-for-your-whole-life pitch dominating the week.
πŸ’‘#35
@ReadFuturist
https://x.com/ReadFuturist/status/2105217020845609389
ReadFuturist flagged Simate, a new Chinese robotics startup founded by Zhang Ying, which topped the global RoboDojo benchmark with its first model, Simate-beta, only three months after launch. The company credits an AI-driven automated research system it calls AutoResearch. If the claim holds, it is one of the first cases of an autoresearch system credited with a frontier result in embodied AI rather than in language models or kernels.
πŸ“‘ Eco Products Radar
Eco Products Radar

GPT-6.1 Sol (9 mentions): the cost baseline in most loop-economics posts, especially for prompt caching.
MCP (7): the boundary layer between loops and tools, now with schema proxies and permission fences around it.
Codex (5) and Claude Code (5): the two harnesses that research loops and software factories are built on.
Jev / TypeSafe (4): small decision models used as gates inside agent loops.
Opus 5.5 (4): the model behind lab-automation and long-running build loops.
Dots (4) and Muse (3): always-on consumer agents whose architectures users are reverse-engineering.
x402 (4): agent payment rails that keep appearing next to loop discussions.
Karpathy's autoresearch (3): still the reference method, now applied to inference, cryptography and even workout plans.
LangChain (3): shipping decision-model middleware inside its agent loop.
← Previous
Super User Daily: 2026-10-02
Next β†’
Ideas Radar: 2026-10-02
← Back to all articles

Comments

Loading...
>_