Loop Daily: July 28, 2026
Autoresearch stopped being a slogan today and turned into plumbing. The headline is AREX, a Chinese deep-research agent whose 4B version reportedly beats a much larger baseline by looping — research, audit every claim, keep what verified, remember what failed, repeat — the clearest picture yet that the first real recursive self-improvement improves the process that builds the next model, not the weights themselves. Underneath it, the conversation got unusually practical: hard retry bounds instead of letting the model decide when to give up, per-key spend caps after someone burned thirty-eight thousand dollars in three hours, memory between loops mattering more than loop speed, and real data showing that turning extended thinking off actually raised pass rates. The most surprising thread is how small the hardware has gotten — a sub-4GB, roughly one-bit model driving itself through a full build-test-fix loop on a used 8GB card. And the applications keep leaving the code editor: photonic design, CFD, structural analysis, quant strategies, even a metro-area social feed cleaned up by an autoresearch loop.
#1
@sudoingX
https://x.com/sudoingX/status/2081453712149721457
A hands-on verdict that Hermes Agent is the strongest autonomous harness right now, based on running the alternatives. The concrete evidence: Hermes drove a tiny local model through a full agentic loop cleanly, in the exact spot where heavier harnesses like OpenClaw could not even find the local endpoint and died at the door. Framed as a firsthand comparison of loop reliability on small local models rather than hype.
https://x.com/sudoingX/status/2081453712149721457
A hands-on verdict that Hermes Agent is the strongest autonomous harness right now, based on running the alternatives. The concrete evidence: Hermes drove a tiny local model through a full agentic loop cleanly, in the exact spot where heavier harnesses like OpenClaw could not even find the local endpoint and died at the door. Framed as a firsthand comparison of loop reliability on small local models rather than hype.
#2
@imjustnewatai
https://x.com/imjustnewatai/status/2081254855167922347
A detailed breakdown of AREX, a Chinese recursively self-improving deep-research agent whose 4B version reportedly beats Qwen3.5-35B on five of six benchmarks. Instead of answering once, AREX researches, audits every claim individually, preserves what it verified, remembers rejected paths, and turns each weak claim into a new targeted research task, looping up to 300 actions and five outer self-improvement rounds while compressing its own context. Ablations credit autonomous context updating (+11.8) and outer recursive auditing (+11.1) for the gains.
https://x.com/imjustnewatai/status/2081254855167922347
A detailed breakdown of AREX, a Chinese recursively self-improving deep-research agent whose 4B version reportedly beats Qwen3.5-35B on five of six benchmarks. Instead of answering once, AREX researches, audits every claim individually, preserves what it verified, remembers rejected paths, and turns each weak claim into a new targeted research task, looping up to 300 actions and five outer self-improvement rounds while compressing its own context. Ablations credit autonomous context updating (+11.8) and outer recursive auditing (+11.1) for the gains.
#3
@sudoingX
https://x.com/sudoingX/status/2081420006244655259
Describes handing Bonsai, a 3.9GB roughly 1-bit model driven by the Hermes agent, a single spec to autonomously build a working rotary (Wankel) engine simulation on an RTX 3060 Ti 8GB. The agent wrote the code, ran its own ten geometry tests, hit failures, read them, and fixed its own lines until the engine spun and all ten tests went green, breaking and recovering three times in a fully observed loop. Presented as a provably correct simulation produced by a tiny local model, not a drawing.
https://x.com/sudoingX/status/2081420006244655259
Describes handing Bonsai, a 3.9GB roughly 1-bit model driven by the Hermes agent, a single spec to autonomously build a working rotary (Wankel) engine simulation on an RTX 3060 Ti 8GB. The agent wrote the code, ran its own ten geometry tests, hit failures, read them, and fixed its own lines until the engine spun and all ten tests went green, breaking and recovering three times in a fully observed loop. Presented as a provably correct simulation produced by a tiny local model, not a drawing.
#4
@imjustnewatai
https://x.com/imjustnewatai/status/2081206432905429421
A grounded essay arguing the first real recursive self-improvement will be unglamorous: a system that improves the process building its successor (training code, kernels, evals, data pipelines, agent scaffolds), which then automates more of the same R&D. It cites concrete primitives already working, the Darwin Godel machine going 20% to 50% on SWE-bench and a self-editing coding agent going 17% to 53%, while noting these improved the scaffold around a frozen model, not the weights. The key threshold is whether each improvement accelerates the next enough to compound faster than compute, data, and safety bottlenecks.
https://x.com/imjustnewatai/status/2081206432905429421
A grounded essay arguing the first real recursive self-improvement will be unglamorous: a system that improves the process building its successor (training code, kernels, evals, data pipelines, agent scaffolds), which then automates more of the same R&D. It cites concrete primitives already working, the Darwin Godel machine going 20% to 50% on SWE-bench and a self-editing coding agent going 17% to 53%, while noting these improved the scaffold around a frozen model, not the weights. The key threshold is whether each improvement accelerates the next enough to compound faster than compute, data, and safety bottlenecks.
#5
@DreggNet
https://x.com/DreggNet/status/2081288406940737553
Reports that their 'Dragon's Egg' autoresearch setup is being unleashed on AWS to find the fastest possible code, under the hard condition that it is not done until it is proven correct against the old implementation and specs. Because they also reason about the compilers, the AI can bypass them directly when it wants. A concrete example of autoresearch driving verified low-level performance optimization rather than open-ended generation.
https://x.com/DreggNet/status/2081288406940737553
Reports that their 'Dragon's Egg' autoresearch setup is being unleashed on AWS to find the fastest possible code, under the hard condition that it is not done until it is proven correct against the old implementation and specs. Because they also reason about the compilers, the AI can bypass them directly when it wants. A concrete example of autoresearch driving verified low-level performance optimization rather than open-ended generation.
#6
@ottogin1
https://x.com/ottogin1/status/2081463918090973470
Reports testing Opus 5 on two real autoresearch tasks, where it delivered surprising results, beating Fable 5 on LLM training tasks and coming only slightly behind Fable 5 on CUDA kernel optimization. A concrete head-to-head benchmark of frontier models inside an autoresearch loop on genuinely hard systems problems.
https://x.com/ottogin1/status/2081463918090973470
Reports testing Opus 5 on two real autoresearch tasks, where it delivered surprising results, beating Fable 5 on LLM training tasks and coming only slightly behind Fable 5 on CUDA kernel optimization. A concrete head-to-head benchmark of frontier models inside an autoresearch loop on genuinely hard systems problems.
#7
@ThePremiseOfIt
https://x.com/ThePremiseOfIt/status/2081437504184459363
A firsthand reliability finding that gpt-5.6-sol consistently and badly regresses their entire training pipeline when run in autoresearch loops, something they never encountered with the other model. They conclude that alignment and consistency, not raw benchmark scores, are what matter for a model you let run unattended in a loop over your codebase.
https://x.com/ThePremiseOfIt/status/2081437504184459363
A firsthand reliability finding that gpt-5.6-sol consistently and badly regresses their entire training pipeline when run in autoresearch loops, something they never encountered with the other model. They conclude that alignment and consistency, not raw benchmark scores, are what matter for a model you let run unattended in a loop over your codebase.
#8
@jackcai1206
https://x.com/jackcai1206/status/2081252992695570674
After extensive hands-on use of many auto-research agents, the user finds it hard to get them to optimize for simple, elegant ideas without a human in the loop. Instead, agents reliably squeeze out gains by increasing complexity, and they can tolerate an enormous amount of it. A sharp practical observation about the failure mode of unsupervised autoresearch: it drifts toward complexity rather than elegance.
https://x.com/jackcai1206/status/2081252992695570674
After extensive hands-on use of many auto-research agents, the user finds it hard to get them to optimize for simple, elegant ideas without a human in the loop. Instead, agents reliably squeeze out gains by increasing complexity, and they can tolerate an enormous amount of it. A sharp practical observation about the failure mode of unsupervised autoresearch: it drifts toward complexity rather than elegance.
#9
@christophcsmith
https://x.com/christophcsmith/status/2081198410158268503
The user had previously built a custom Bluesky feed for their metro area that just regex-matched city names and produced many false positives. They ran an auto-research loop over it and got it into much better shape. A concrete, non-frontier everyday use of an autoresearch loop to iteratively improve a content-filtering feed.
https://x.com/christophcsmith/status/2081198410158268503
The user had previously built a custom Bluesky feed for their metro area that just regex-matched city names and produced many false positives. They ran an auto-research loop over it and got it into much better shape. A concrete, non-frontier everyday use of an autoresearch loop to iteratively improve a content-filtering feed.
#10
@luoyu_toc
https://x.com/luoyu_toc/status/2081256558252662822
An EMNLP emergency reviewer's reflection after spending five hours on one paper: the motivation was building a rocket while the method was tightening a screw. They observe that Auto Research is making it increasingly easy to generate exciting-sounding motivations, but turning that into a convincing method and solid evidence remains one of the hardest and least automatable parts of research. A pointed critique of where autoresearch currently helps and where it does not.
https://x.com/luoyu_toc/status/2081256558252662822
An EMNLP emergency reviewer's reflection after spending five hours on one paper: the motivation was building a rocket while the method was tightening a screw. They observe that Auto Research is making it increasingly easy to generate exciting-sounding motivations, but turning that into a convincing method and solid evidence remains one of the hardest and least automatable parts of research. A pointed critique of where autoresearch currently helps and where it does not.
#11
@SasuRobert
https://x.com/SasuRobert/status/2081429156147974522
Reports that flash-lite is the model they use to benchmark and autoresearch their agentic systems, because it is the easiest to build with via batch API and to test complex loops, graphs, evals, and RL techniques. The very fast responses let their autoresearch scale quickly. A concrete workflow note on picking a cheap fast model as the workhorse for autoresearch iteration.
https://x.com/SasuRobert/status/2081429156147974522
Reports that flash-lite is the model they use to benchmark and autoresearch their agentic systems, because it is the easiest to build with via batch API and to test complex loops, graphs, evals, and RL techniques. The very fast responses let their autoresearch scale quickly. A concrete workflow note on picking a cheap fast model as the workhorse for autoresearch iteration.
#12
@FakePsyho
https://x.com/FakePsyho/status/2081519777420300392
A benchmarking veteran explains why good autoresearch evals barely exist: testing interactive problems or autoresearch capabilities with a decent sample size is extremely expensive, on the order of $10,000+ per single eval, and no one wants to pay it. They add that most benchmarks they have examined are either slop or test a very narrow capability. A concrete diagnosis of the measurement gap holding autoresearch back.
https://x.com/FakePsyho/status/2081519777420300392
A benchmarking veteran explains why good autoresearch evals barely exist: testing interactive problems or autoresearch capabilities with a decent sample size is extremely expensive, on the order of $10,000+ per single eval, and no one wants to pay it. They add that most benchmarks they have examined are either slop or test a very narrow capability. A concrete diagnosis of the measurement gap holding autoresearch back.
#13
@ariccio
https://x.com/ariccio/status/2081511612226085201
The user says they are waiting for someone to run a frontier model in an autoresearch loop with CFD, and would do it themselves but their tokens are already assigned to similar engineering challenges: truss/structural analysis, two-pipe steam-system modeling, and golf analytics for their dad. A concrete glimpse of autoresearch being applied to real physical-engineering problems rather than coding.
https://x.com/ariccio/status/2081511612226085201
The user says they are waiting for someone to run a frontier model in an autoresearch loop with CFD, and would do it themselves but their tokens are already assigned to similar engineering challenges: truss/structural analysis, two-pipe steam-system modeling, and golf analytics for their dad. A concrete glimpse of autoresearch being applied to real physical-engineering problems rather than coding.
#14
@minhash
https://x.com/minhash/status/2081312620783784095
Argues autoresearch is potentially the single most important outcome of AI, especially for domains with verifiable rewards like AI, systems, hardware, engineering, and medicine, yet it feels severely underdone right now. They ask why more companies are not focused on doing it better. A concise market take flagging autoresearch infrastructure as an underbuilt opportunity.
https://x.com/minhash/status/2081312620783784095
Argues autoresearch is potentially the single most important outcome of AI, especially for domains with verifiable rewards like AI, systems, hardware, engineering, and medicine, yet it feels severely underdone right now. They ask why more companies are not focused on doing it better. A concise market take flagging autoresearch infrastructure as an underbuilt opportunity.
#15
@LeeLeepenkman
https://x.com/LeeLeepenkman/status/2081243329065550330
Reports pointing their 'run forever' agent at nearly anything and having it keep improving the target until it can't, drifting onto an auto-research improvement curve that gets stronger with stronger base models and stronger still with a human in the loop. They tie this to earlier Ralph-loop experiments and argue recursive improvement quietly emerged out of long-running coding agents well before it had a name.
https://x.com/LeeLeepenkman/status/2081243329065550330
Reports pointing their 'run forever' agent at nearly anything and having it keep improving the target until it can't, drifting onto an auto-research improvement curve that gets stronger with stronger base models and stronger still with a human in the loop. They tie this to earlier Ralph-loop experiments and argue recursive improvement quietly emerged out of long-running coding agents well before it had a name.
#16
@twerpzz
https://x.com/twerpzz/status/2081407634205241638
Describes building a self-improving loop at work based on an ecc auto-research setup, where their sessions get graded and their skills, troubleshooting, and token use measurably improve over time because everything is persisted in GitHub. A concrete personal example of an autoresearch-style feedback loop applied to improving a human operator's own workflow, not just a model.
https://x.com/twerpzz/status/2081407634205241638
Describes building a self-improving loop at work based on an ecc auto-research setup, where their sessions get graded and their skills, troubleshooting, and token use measurably improve over time because everything is persisted in GitHub. A concrete personal example of an autoresearch-style feedback loop applied to improving a human operator's own workflow, not just a model.
#17
@quantum_jake
https://x.com/quantum_jake/status/2081480883731873852
Points to a talk from Flexcompute on how they are using auto research for photonic design. A concrete pointer to autoresearch being applied in a specialized hardware-design domain (photonics) with a real company behind it.
https://x.com/quantum_jake/status/2081480883731873852
Points to a talk from Flexcompute on how they are using auto research for photonic design. A concrete pointer to autoresearch being applied in a specialized hardware-design domain (photonics) with a real company behind it.
#18
@trashpandaemoji
https://x.com/trashpandaemoji/status/2081467064662049034
Proposes that instead of building the same thing multiple times, agents should continuously fork at high-impact decision points, then collapse the branch once enough is learned and incorporate the learnings. They frame automating with agents as basically auto-research applied to building software. A methodology note on branch-and-collapse exploration for agentic building.
https://x.com/trashpandaemoji/status/2081467064662049034
Proposes that instead of building the same thing multiple times, agents should continuously fork at high-impact decision points, then collapse the branch once enough is learned and incorporate the learnings. They frame automating with agents as basically auto-research applied to building software. A methodology note on branch-and-collapse exploration for agentic building.
#19
@silentguyy66
https://x.com/silentguyy66/status/2081311870862979410
Amplifies Karpathy's point that sophisticated memory is not yet implemented in agents, it is just context compaction when the window runs out, so the next iteration starts blind. They argue everyone optimizes loop speed while almost nobody builds persistent context between loops, and that a memory layer between loops matters more than loop speed for compounding. A clear framing of the memory-between-loops bottleneck.
https://x.com/silentguyy66/status/2081311870862979410
Amplifies Karpathy's point that sophisticated memory is not yet implemented in agents, it is just context compaction when the window runs out, so the next iteration starts blind. They argue everyone optimizes loop speed while almost nobody builds persistent context between loops, and that a memory layer between loops matters more than loop speed for compounding. A clear framing of the memory-between-loops bottleneck.
#20
@VibeCoderOfek
https://x.com/VibeCoderOfek/status/2081259867159810222
Argues self-improving harnesses are the real systems problem now, because most agent work still treats the loop as a black box instead of something that can rewrite its own scaffolding under hard constraints. A concise thesis that the frontier of agent engineering is a loop that edits its own harness.
https://x.com/VibeCoderOfek/status/2081259867159810222
Argues self-improving harnesses are the real systems problem now, because most agent work still treats the loop as a black box instead of something that can rewrite its own scaffolding under hard constraints. A concise thesis that the frontier of agent engineering is a loop that edits its own harness.
#21
@VontariusF
https://x.com/VontariusF/status/2081251903594188805
Announces open-sourcing a CLI that plugs into any agent system and does government-contract discovery plus grant sourcing and writing based on your company, on autopilot. It is local-first, accurate, and self-improving when used with Hermes. A concrete self-improving agent tool aimed at a specific non-coding business workflow.
https://x.com/VontariusF/status/2081251903594188805
Announces open-sourcing a CLI that plugs into any agent system and does government-contract discovery plus grant sourcing and writing based on your company, on autopilot. It is local-first, accurate, and self-improving when used with Hermes. A concrete self-improving agent tool aimed at a specific non-coding business workflow.
#22
@DrElectronX
https://x.com/DrElectronX/status/2081353018436641213
Sketches a '/jarvis mode' agent that sets up git, policies, SOPs, organization, and a hierarchical searchable memory system, installs best practices, and self-improves the codebase and issues in a prioritized manner without bothering the user unless truly needed, guiding them with questions and suggestions. A concrete spec for a self-directed, self-improving operator agent.
https://x.com/DrElectronX/status/2081353018436641213
Sketches a '/jarvis mode' agent that sets up git, policies, SOPs, organization, and a hierarchical searchable memory system, installs best practices, and self-improves the codebase and issues in a prioritized manner without bothering the user unless truly needed, guiding them with questions and suggestions. A concrete spec for a self-directed, self-improving operator agent.
#23
@echantech1
https://x.com/echantech1/status/2081299231848169830
Shares honest lessons from building an agentic loop: they spent about 20 minutes going back and forth writing the loop markdown and script just to tune it, defining success/failure/acceptance criteria matters a lot, and loops should target exactly one bottleneck. They candidly admit that using loops to attack bottlenecks is often just being lazy, creating a 'slop cannon' to identify and prove root cause. A grounded account of the up-front investment a real loop demands.
https://x.com/echantech1/status/2081299231848169830
Shares honest lessons from building an agentic loop: they spent about 20 minutes going back and forth writing the loop markdown and script just to tune it, defining success/failure/acceptance criteria matters a lot, and loops should target exactly one bottleneck. They candidly admit that using loops to attack bottlenecks is often just being lazy, creating a 'slop cannon' to identify and prove root cause. A grounded account of the up-front investment a real loop demands.
#24
@RonnyBruknapp
https://x.com/RonnyBruknapp/status/2081399978598056241
Explains that when the same model performs differently across Claude Code and Cursor, it is the harness, not the model: Claude Code gives it more of the repo and a real agentic loop, while Cursor keeps you in tight inline edits. They argue 'X beats Y' usually just means you prefer that workflow, not that the model improved. A clear articulation of harness-as-variable.
https://x.com/RonnyBruknapp/status/2081399978598056241
Explains that when the same model performs differently across Claude Code and Cursor, it is the harness, not the model: Claude Code gives it more of the repo and a real agentic loop, while Cursor keeps you in tight inline edits. They argue 'X beats Y' usually just means you prefer that workflow, not that the model improved. A clear articulation of harness-as-variable.
#25
@faisalusuf
https://x.com/faisalusuf/status/2081172350272463224
A cautionary firsthand note that a black-box agentic loop with no control is why some organizations got surprise bills of millions of dollars from Anthropic. They recommend being cautious and ready for surprises when handing work to an uncontrolled loop. A concrete warning about the cost risk of unbounded autonomous loops.
https://x.com/faisalusuf/status/2081172350272463224
A cautionary firsthand note that a black-box agentic loop with no control is why some organizations got surprise bills of millions of dollars from Anthropic. They recommend being cautious and ready for surprises when handing work to an uncontrolled loop. A concrete warning about the cost risk of unbounded autonomous loops.
#26
@MoezZhioua
https://x.com/MoezZhioua/status/2081407045677298049
Describes building a full agent loop in just 792 lines with only four tools, arguing that skipping MCP frees tens of thousands of tokens and tightens the context budget. The takeaway is to keep the harness minimal so the token budget wins, rather than loading heavy tool schemas the model rarely needs.
https://x.com/MoezZhioua/status/2081407045677298049
Describes building a full agent loop in just 792 lines with only four tools, arguing that skipping MCP frees tens of thousands of tokens and tightens the context budget. The takeaway is to keep the harness minimal so the token budget wins, rather than loading heavy tool schemas the model rarely needs.
#27
@denogrowth
https://x.com/denogrowth/status/2081359395347128440
Recounts a case where one employee racked up $38,000 in AI charges in three hours because no per-key spend cap was set, warning that a single agent loop can outspend a whole team's monthly budget in an afternoon and that the cost only surfaces after the bill lands. Urges builders to set hard caps, alerts, and per-key limits before handing out live keys.
https://x.com/denogrowth/status/2081359395347128440
Recounts a case where one employee racked up $38,000 in AI charges in three hours because no per-key spend cap was set, warning that a single agent loop can outspend a whole team's monthly budget in an afternoon and that the cost only surfaces after the bill lands. Urges builders to set hard caps, alerts, and per-key limits before handing out live keys.
#28
@ja818_
https://x.com/ja818_/status/2081354345707348226
Describes building Houston, an agent system that deleted its bundled CLI provider binaries (now running providers in-process) and removed its MCP client entirely, keeping only one in-process MCP server bridging five tools to the backend. They argue the real axis is not CLI vs MCP but what the model sees in context and who holds the credential, and expose 1000+ integrations via just two tools (integration_search and integration_execute) rather than thousands of schemas.
https://x.com/ja818_/status/2081354345707348226
Describes building Houston, an agent system that deleted its bundled CLI provider binaries (now running providers in-process) and removed its MCP client entirely, keeping only one in-process MCP server bridging five tools to the backend. They argue the real axis is not CLI vs MCP but what the model sees in context and who holds the credential, and expose 1000+ integrations via just two tools (integration_search and integration_execute) rather than thousands of schemas.
#29
@stretchcloud
https://x.com/stretchcloud/status/2081512135167988068
Summarizes a Microsoft telemetry study of tens of thousands of engineers finding Copilot CLI users merged 24.9% more PRs (50.1% with five-plus days of use), a 2.2x lift over Claude Code users in the same population. They argue Copilot's tighter GitHub/PR-workflow integration beats Claude Code's broader agent loop for PR output, and that the bottleneck is workflow integration, not model capability.
https://x.com/stretchcloud/status/2081512135167988068
Summarizes a Microsoft telemetry study of tens of thousands of engineers finding Copilot CLI users merged 24.9% more PRs (50.1% with five-plus days of use), a 2.2x lift over Claude Code users in the same population. They argue Copilot's tighter GitHub/PR-workflow integration beats Claude Code's broader agent loop for PR output, and that the bottleneck is workflow integration, not model capability.
#30
@DogukanUrker
https://x.com/DogukanUrker/status/2081278089212948612
Diagnoses that an agent loop losing its tool definitions mid-run and making bad tool calls is likely Ollama's default num_ctx truncating context to 4096 tokens, not a model failure. Recommends retrying with num_ctx bumped, or running straight through llama.cpp, which is how they ran it. A concrete, actionable fix for a common local-agent loop breakdown.
https://x.com/DogukanUrker/status/2081278089212948612
Diagnoses that an agent loop losing its tool definitions mid-run and making bad tool calls is likely Ollama's default num_ctx truncating context to 4096 tokens, not a model failure. Recommends retrying with num_ctx bumped, or running straight through llama.cpp, which is how they ran it. A concrete, actionable fix for a common local-agent loop breakdown.
#31
@MichaelGannotti
https://x.com/MichaelGannotti/status/2081449488041062604
Reports an 'offlabel guide' finding that enabling extended thinking is net-negative on held-out work: it fabricates bugs in clean code, over-refuses authorized work, and once hung an agent loop at turn 11 for 91 minutes. They cite concrete pass rates of 94.2% with thinking off versus 91.3% with thinking on. A data-backed argument against reflexively maxing reasoning effort inside loops.
https://x.com/MichaelGannotti/status/2081449488041062604
Reports an 'offlabel guide' finding that enabling extended thinking is net-negative on held-out work: it fabricates bugs in clean code, over-refuses authorized work, and once hung an agent loop at turn 11 for 91 minutes. They cite concrete pass rates of 94.2% with thinking off versus 91.3% with thinking on. A data-backed argument against reflexively maxing reasoning effort inside loops.
#32
@anton_bt
https://x.com/anton_bt/status/2081390359205040628
Cleanly separates three architectures people conflate: a harness (the controlled environment of tools, permissions, memory, budgets, retries), a loop (PLAN to ACT to VERIFY to REPAIR, driven by structured feedback from a real verifier), and a graph (branching workflows with routes and audit trails). They stress the value of a loop is structured feedback, not blind retry, and prescribe building the harness first, adding a loop once you have a verifier, and only adding a graph when the workflow genuinely branches.
https://x.com/anton_bt/status/2081390359205040628
Cleanly separates three architectures people conflate: a harness (the controlled environment of tools, permissions, memory, budgets, retries), a loop (PLAN to ACT to VERIFY to REPAIR, driven by structured feedback from a real verifier), and a graph (branching workflows with routes and audit trails). They stress the value of a loop is structured feedback, not blind retry, and prescribe building the harness first, adding a loop once you have a verifier, and only adding a graph when the workflow genuinely branches.
#33
@ZayahNelson0
https://x.com/ZayahNelson0/status/2081190505145331860
Describes running a daily agent loop for job applications plus cold email that operates a real browser and routes across 43 free-tier API keys spread over 14 providers, running at $0 per month. A concrete example of a fully autonomous personal loop whose main engineering trick is aggressive multi-provider key rotation to stay free.
https://x.com/ZayahNelson0/status/2081190505145331860
Describes running a daily agent loop for job applications plus cold email that operates a real browser and routes across 43 free-tier API keys spread over 14 providers, running at $0 per month. A concrete example of a fully autonomous personal loop whose main engineering trick is aggressive multi-provider key rotation to stay free.
#34
@stretchcloud
https://x.com/stretchcloud/status/2081204105515811180
Explains an architecture decision to pick turbopuffer over pgvector and others for the agentic memory retrieval tier: it runs entirely on object storage (~70x cheaper than RAM+SSD), writes ~1,190 chunks/sec versus pgvector's 540-760, and powers Cursor indexing over a trillion chunks across 80 million namespaces. The argument is that the hidden bottleneck in agentic memory is cost-per-namespace as agents proliferate, not recall quality.
https://x.com/stretchcloud/status/2081204105515811180
Explains an architecture decision to pick turbopuffer over pgvector and others for the agentic memory retrieval tier: it runs entirely on object storage (~70x cheaper than RAM+SSD), writes ~1,190 chunks/sec versus pgvector's 540-760, and powers Cursor indexing over a trillion chunks across 80 million namespaces. The argument is that the hidden bottleneck in agentic memory is cost-per-namespace as agents proliferate, not recall quality.
#35
@EddyWoodss
https://x.com/EddyWoodss/status/2081247870154330287
Announces shipping a free 'agent loop detector' tool that catches recursive self-improvement spirals before they burn through credits, released alongside two other tools. A concrete piece of loop-safety tooling aimed at the exact runaway-cost failure mode others in this space keep warning about.
https://x.com/EddyWoodss/status/2081247870154330287
Announces shipping a free 'agent loop detector' tool that catches recursive self-improvement spirals before they burn through credits, released alongside two other tools. A concrete piece of loop-safety tooling aimed at the exact runaway-cost failure mode others in this space keep warning about.
#36
@neil_xbt
https://x.com/neil_xbt/status/2081231787616076219
Reports that across 70 studied real-world agent-loop implementations, a large share had zero formal bounds on how many times a failed step would be retried, so failure behavior was whatever the model decided in the moment. Advocates strict escalation: a defined number of recovery attempts, a defined success criterion, and a hard handoff to a human the instant the limit is hit, with no creative fifth attempt at the same failed approach.
https://x.com/neil_xbt/status/2081231787616076219
Reports that across 70 studied real-world agent-loop implementations, a large share had zero formal bounds on how many times a failed step would be retried, so failure behavior was whatever the model decided in the moment. Advocates strict escalation: a defined number of recovery attempts, a defined success criterion, and a hard handoff to a human the instant the limit is hit, with no creative fifth attempt at the same failed approach.
#37
@mernit
https://x.com/mernit/status/2081436968219800026
Shares agent-skill tips: agents perform much better with human-written skills than AI-generated ones, and a good skill specifies exactly three things, the inputs, the outputs, and what 'good' looks like. They recommend iterating by running the agent loop and updating those three fields based on each run until the skill behaves as wanted.
https://x.com/mernit/status/2081436968219800026
Shares agent-skill tips: agents perform much better with human-written skills than AI-generated ones, and a good skill specifies exactly three things, the inputs, the outputs, and what 'good' looks like. They recommend iterating by running the agent loop and updating those three fields based on each run until the skill behaves as wanted.
#38
@liocoh
https://x.com/liocoh/status/2081434944144801999
Flags a genuinely under-covered topic: running open models as unattended agents. Everyone benchmarks chat quality, but almost no one shows what happens when a local model runs an agent loop for hours, or how to catch it failing before the errors compound. A pointed observation about the missing evaluation regime for long-horizon local agents.
https://x.com/liocoh/status/2081434944144801999
Flags a genuinely under-covered topic: running open models as unattended agents. Everyone benchmarks chat quality, but almost no one shows what happens when a local model runs an agent loop for hours, or how to catch it failing before the errors compound. A pointed observation about the missing evaluation regime for long-horizon local agents.
📡 Eco Products Radar
Eco Products Radar
MCP — the tooling layer everyone is re-litigating; several argue skipping or minimizing it frees tens of thousands of tokens per loop.
Hermes Agent — the harness of choice for driving tiny local models through full build-test-fix loops with persistent memory.
Codex — the recurring open, auditable-loop alternative people run alongside or against Claude Code.
AREX — the week's most-discussed project: an open recursively self-improving deep-research agent shipped in 4B and 122B sizes.
Bonsai — a roughly 3.9GB, one-bit model that keeps surprising people by holding an agent loop past turn three on cheap hardware.
MCP — the tooling layer everyone is re-litigating; several argue skipping or minimizing it frees tens of thousands of tokens per loop.
Hermes Agent — the harness of choice for driving tiny local models through full build-test-fix loops with persistent memory.
Codex — the recurring open, auditable-loop alternative people run alongside or against Claude Code.
AREX — the week's most-discussed project: an open recursively self-improving deep-research agent shipped in 4B and 122B sizes.
Bonsai — a roughly 3.9GB, one-bit model that keeps surprising people by holding an agent loop past turn three on cheap hardware.
Comments