Loop Daily: 2026-10-11
The loop went decentralized this window. One operator published the overnight ledger of five Fable 5.1 agents orchestrating 92 sandbox workers over the OpenAI math dump: more than 700 million tokens, five new upper bounds for integer multiplication, the Sidorenko residual cut from 978 to 189, a new Van der Waerden coefficient, and a leaderboard where some of the top contributors are not mathematicians but people who are good at making agents think faster. The same thread reported that every structural jump near the current wall came from Fable while Codex teams owned the fast composition work, and that a Lean formalization arrived so mathematicians could trust and join. Elsewhere Bend's author laid out a plan to let the compiler evolve itself under autoresearch plus laws, SenseNova doubled a robotics score through harness-level self-improvement without touching weights, Meta's IdeaScientist beat Claude Code SDK and Codex SDK on research novelty with a 27B open model, a 12 GB RTX 3060 built a playable game over five hours through Hermes, and Microsoft trained a coding agent inside its real harness. The practitioner notes converge on the same lesson: strip tools down to primitives, define done before the model touches anything, give every loop a budget, and record the decisions, not just the diffs.
#1
@RohanArun
https://x.com/RohanArun/status/2108910833983979590
RohanArun compiled the stats from running a fleet of agents overnight on the OpenAI math dump: five Fable 5.1 agents orchestrating 92 sandbox workers spent more than 700 million tokens and came back with five new upper bounds for integer multiplication, a Unit Distances grid shrunk to 12 edges, the Sidorenko residual constant cut from 978 to 189 and a new Van der Waerden coefficient of 1/25. Three observations: more agents does not accelerate solutions linearly, so human researchers with good intuition are needed to guide independent fleets; some of the top contributors are not math experts but experts at using agents to think faster; and most results were one-shot, with only integer multiplication needing a fleet of sandboxes to test new regimes. A call with the top contributors is set to share best practices, and the author is recruiting agent wranglers and infra sponsors.
https://x.com/RohanArun/status/2108910833983979590
RohanArun compiled the stats from running a fleet of agents overnight on the OpenAI math dump: five Fable 5.1 agents orchestrating 92 sandbox workers spent more than 700 million tokens and came back with five new upper bounds for integer multiplication, a Unit Distances grid shrunk to 12 edges, the Sidorenko residual constant cut from 978 to 189 and a new Van der Waerden coefficient of 1/25. Three observations: more agents does not accelerate solutions linearly, so human researchers with good intuition are needed to guide independent fleets; some of the top contributors are not math experts but experts at using agents to think faster; and most results were one-shot, with only integer multiplication needing a fleet of sandboxes to test new regimes. A call with the top contributors is set to share best practices, and the author is recruiting agent wranglers and infra sponsors.
#2
@RohanArun
https://x.com/RohanArun/status/2108943905328615654
RohanArun's progress report on the integer multiplication bound, which moved 6 percent in 24 hours: four new bit words in a day accounted for about 60 percent of the gain, and when the bit side came within 0.6 percent of the complex side's cap the community re-annealed the complex modules within minutes and raised the cap. Two contributions went into the repo core: a Lean formalization verifying key finite identities and exact arithmetic, so mathematicians who trust Lean can contribute in parallel, and scheduling as a composable stage. The split by model is the striking part: as the wall approached, every big leap started coming from Fable 5.1 rather than Astra on high, every structural jump since the previous day was disclosed as Claude-assisted, while Codex-assisted teams owned the fast composition and admission work. The best breakthroughs still come from a human pointing the model at a different object to force out-of-distribution solutions, and the author calls what is forming a new kind of decentralized auto-research that moves faster than papers can be published.
https://x.com/RohanArun/status/2108943905328615654
RohanArun's progress report on the integer multiplication bound, which moved 6 percent in 24 hours: four new bit words in a day accounted for about 60 percent of the gain, and when the bit side came within 0.6 percent of the complex side's cap the community re-annealed the complex modules within minutes and raised the cap. Two contributions went into the repo core: a Lean formalization verifying key finite identities and exact arithmetic, so mathematicians who trust Lean can contribute in parallel, and scheduling as a composable stage. The split by model is the striking part: as the wall approached, every big leap started coming from Fable 5.1 rather than Astra on high, every structural jump since the previous day was disclosed as Claude-assisted, while Codex-assisted teams owned the fast composition and admission work. The best breakthroughs still come from a human pointing the model at a different object to force out-of-distribution solutions, and the author calls what is forming a new kind of decentralized auto-research that moves faster than papers can be published.
#3
@RohanArun
https://x.com/RohanArun/status/2108789458145259574
RohanArun on the daily rhythm of this work: ideas get queued up for agents that run overnight around the clock, and after a few months of doing auto-research this way the author frequently wakes up to a breakthrough. That night's result was a new conditional bound submitted with Fable 5.1 delivering the final step.
https://x.com/RohanArun/status/2108789458145259574
RohanArun on the daily rhythm of this work: ideas get queued up for agents that run overnight around the clock, and after a few months of doing auto-research this way the author frequently wakes up to a breakthrough. That night's result was a new conditional bound submitted with Fable 5.1 delivering the final step.
#4
@VictorTaelin
https://x.com/VictorTaelin/status/2108868998876033468
VictorTaelin explained the under-told half of Bend's thesis: the language was designed to scale with AI, and the goal is for Bend to start writing itself, with agents and laws, improving at an unprecedented rate. The V1 is in TypeScript, which has no laws, so the kernel was kept tiny at around 100K tokens so the whole thing fits in a model's context for a full port, after which laws are set up and the compiler evolves itself through autoresearch plus laws. The open problem is getting agents to write fast code, not just correct code, and the plan is a formalized CUDA model in Bend so agents can prove statements about executable performance, which doubles as a way to prove the compiler correct. The author expects to start around December and, in a follow-up, argues Bend is the best candidate for the laws-plus-autoresearch loop because a certified compiler has a clean law: it is right when it matches the interpreter.
https://x.com/VictorTaelin/status/2108868998876033468
VictorTaelin explained the under-told half of Bend's thesis: the language was designed to scale with AI, and the goal is for Bend to start writing itself, with agents and laws, improving at an unprecedented rate. The V1 is in TypeScript, which has no laws, so the kernel was kept tiny at around 100K tokens so the whole thing fits in a model's context for a full port, after which laws are set up and the compiler evolves itself through autoresearch plus laws. The open problem is getting agents to write fast code, not just correct code, and the plan is a formalized CUDA model in Bend so agents can prove statements about executable performance, which doubles as a way to prove the compiler correct. The author expects to start around December and, in a follow-up, argues Bend is the best candidate for the laws-plus-autoresearch loop because a certified compiler has a clean law: it is right when it matches the interpreter.
#5
@nrqa__
https://x.com/nrqa__/status/2108934827961696710
nrqa__ relayed that SenseNova-RoboRSI nearly doubled a RoboDojo score with the same model by making the harness self-improving instead of training a larger model: 56.83 average against the public GPT-6 Astra baseline of 28.97, a 96.2 percent relative gain with no change to the base weights, plus 94.50 on LIBERO-PRO and 64.60 on RoboCasa. The loop analyzes execution traces, finds weaknesses, explores alternative strategies, evaluates repeatedly and folds successful changes into future versions of the agent, optimizing how it plans, calls tools, executes and learns from physical feedback. The repository is live with code and a technical report to follow.
https://x.com/nrqa__/status/2108934827961696710
nrqa__ relayed that SenseNova-RoboRSI nearly doubled a RoboDojo score with the same model by making the harness self-improving instead of training a larger model: 56.83 average against the public GPT-6 Astra baseline of 28.97, a 96.2 percent relative gain with no change to the base weights, plus 94.50 on LIBERO-PRO and 64.60 on RoboCasa. The loop analyzes execution traces, finds weaknesses, explores alternative strategies, evaluates repeatedly and folds successful changes into future versions of the agent, optimizing how it plans, calls tools, executes and learns from physical feedback. The repository is live with code and a technical report to follow.
#6
@dair_ai
https://x.com/dair_ai/status/2108935723105841430
dair_ai flagged a Meta paper on AI research agents: IdeaScientist, a 27B open model that beats the strongest open autoresearch baseline by 14.0 percent, mostly on novelty, and beats Claude Code SDK and Codex SDK setups by up to 5.9 percent. Ideation is split into three roles trained separately with RL: a gap finder that reads related work for limitations, an innovator that retrieves mechanisms that solved similar problems in other fields, and a writer that turns the result into a full proposal, with retrieval over a corpus of 2.77 million decomposed research ideas. Evaluation only allows literature before a cutoff and scores proposals against directions later explored in 15,000 human-written papers. A reply asked whether the eval set was held out from the autoresearch training loop, and another asked how novelty was judged and whether the gains survive human review.
https://x.com/dair_ai/status/2108935723105841430
dair_ai flagged a Meta paper on AI research agents: IdeaScientist, a 27B open model that beats the strongest open autoresearch baseline by 14.0 percent, mostly on novelty, and beats Claude Code SDK and Codex SDK setups by up to 5.9 percent. Ideation is split into three roles trained separately with RL: a gap finder that reads related work for limitations, an innovator that retrieves mechanisms that solved similar problems in other fields, and a writer that turns the result into a full proposal, with retrieval over a corpus of 2.77 million decomposed research ideas. Evaluation only allows literature before a cutoff and scores proposals against directions later explored in 15,000 human-written papers. A reply asked whether the eval set was held out from the autoresearch training loop, and another asked how novelty was judged and whether the gains survive human review.
#7
@Oluwaphilemon1
https://x.com/Oluwaphilemon1/status/2108401318540779749
Oluwaphilemon1 documented a 12 GB RTX 3060 that spent five hours building an entire playable game with no handwritten code: Bonsai 2, a ternary-compressed Qwen3.8-27B squeezed into a 5.95 GB weight file, running through Hermes Agent with a 125K context at about 50 tokens per second fresh and 22 on average. By the end the agent had generated 328K tokens, eight JavaScript files and 2,368 lines. The point the author draws is that the agent loop mattered more than raw token speed: creating files, holding project context, revisiting earlier code and recovering from errors over hours is a different test from an impressive first response. Speculative decoding roughly doubled fresh decode speed from 26 to 50 tokens per second, and the caveat is that a successful build does not prove the game is bug-free.
https://x.com/Oluwaphilemon1/status/2108401318540779749
Oluwaphilemon1 documented a 12 GB RTX 3060 that spent five hours building an entire playable game with no handwritten code: Bonsai 2, a ternary-compressed Qwen3.8-27B squeezed into a 5.95 GB weight file, running through Hermes Agent with a 125K context at about 50 tokens per second fresh and 22 on average. By the end the agent had generated 328K tokens, eight JavaScript files and 2,368 lines. The point the author draws is that the agent loop mattered more than raw token speed: creating files, holding project context, revisiting earlier code and recovering from errors over hours is a different test from an impressive first response. Speculative decoding roughly doubled fresh decode speed from 26 to 50 tokens per second, and the caveat is that a successful build does not prove the game is bug-free.
#8
@MikeTamir
https://x.com/MikeTamir/status/2108937857553351136
MikeTamir shared Agora, which uses Git as a shared memory DAG for decentralized multi-agent autoresearch, letting 13 LLM workers collaboratively solve weight-transfer tasks without central planning. A second post on the same work frames the branchable Git repository as persistent memory and asks whether it becomes the next engine for auto-research and scientific discovery.
https://x.com/MikeTamir/status/2108937857553351136
MikeTamir shared Agora, which uses Git as a shared memory DAG for decentralized multi-agent autoresearch, letting 13 LLM workers collaboratively solve weight-transfer tasks without central planning. A second post on the same work frames the branchable Git repository as persistent memory and asks whether it becomes the next engine for auto-research and scientific discovery.
#9
@Ryun1_
https://x.com/Ryun1_/status/2108587212933591524
Ryun1_ reports that autoresearch challenges produce powerful results, such as 5x smaller STARK proofs, but work is often siloed between participants and effort overlaps. The project being built lets agents work collaboratively toward a shared landscape of knowledge: they discuss, experiment and fail in public, sharing their notes, on the bet that this kind of collaboration gets further than parallel isolated runs.
https://x.com/Ryun1_/status/2108587212933591524
Ryun1_ reports that autoresearch challenges produce powerful results, such as 5x smaller STARK proofs, but work is often siloed between participants and effort overlaps. The project being built lets agents work collaboratively toward a shared landscape of knowledge: they discuss, experiment and fail in public, sharing their notes, on the bet that this kind of collaboration gets further than parallel isolated runs.
#10
@KasparsDancis
https://x.com/KasparsDancis/status/2108804321748148696
KasparsDancis is running Pi-autoresearch overnight to see whether it can improve Whimsical's app performance test metrics by optimizing the Clojurific core and compiler while maintaining correctness. A production app's own performance suite is the metric, and the compiler is the thing being edited.
https://x.com/KasparsDancis/status/2108804321748148696
KasparsDancis is running Pi-autoresearch overnight to see whether it can improve Whimsical's app performance test metrics by optimizing the Clojurific core and compiler while maintaining correctness. A production app's own performance suite is the metric, and the compiler is the thing being edited.
#11
@Shekswess
https://x.com/Shekswess/status/2109017519172276346
Shekswess is doing DAG-based auto-research with Codex, and the part that impressed the author is Codex using computer-use to open the web UI and inspect the data itself to analyze results, rather than only reading numbers back from the terminal.
https://x.com/Shekswess/status/2109017519172276346
Shekswess is doing DAG-based auto-research with Codex, and the part that impressed the author is Codex using computer-use to open the web UI and inspect the data itself to analyze results, rather than only reading numbers back from the terminal.
#12
@al3rez
https://x.com/al3rez/status/2108961002465394892
al3rez summarized Microsoft's Agent Lightning v1.0: a coding agent trained inside its real harness took Qwen3.5-9B from 41.8 to 56.4 percent on SWE-bench Verified with around 6,000 training samples, in roughly 3,500 lines of code. Most agent RL setups force developers to rebuild the agent loop inside a training framework, so the model learns in one environment and runs in another; Agent Lightning instead puts a proxy between the model and the existing harness, which keeps its tools, context management and execution loop while the trainer collects trajectories for RL, and runs agents as Kubernetes jobs on your own infrastructure. The author's takeaway: train the model where it will actually work.
https://x.com/al3rez/status/2108961002465394892
al3rez summarized Microsoft's Agent Lightning v1.0: a coding agent trained inside its real harness took Qwen3.5-9B from 41.8 to 56.4 percent on SWE-bench Verified with around 6,000 training samples, in roughly 3,500 lines of code. Most agent RL setups force developers to rebuild the agent loop inside a training framework, so the model learns in one environment and runs in another; Agent Lightning instead puts a proxy between the model and the existing harness, which keeps its tools, context management and execution loop while the trainer collects trajectories for RL, and runs agents as Kubernetes jobs on your own infrastructure. The author's takeaway: train the model where it will actually work.
#13
@Ankur_Samanta_
https://x.com/Ankur_Samanta_/status/2108634794271637830
Ankur_Samanta_ on MIRA-AC: models can learn effectively from their own proxy hill-climbing signals, improving gold outcomes across several autoresearch environments and transferring to one held out from policy training. The method uses decision-level credit assignment, so optimization focuses on the decisions that steer the direction of research rather than every token, about 25 percent of the decision and execution tokens in an episode. Training uses buffered asynchronous single-rollout policy optimization because research episodes vary widely in length: finished slots fill with new questions immediately, completed episodes feed decoupled actor and critic buffers, and live episodes pick up new weights at their next decision.
https://x.com/Ankur_Samanta_/status/2108634794271637830
Ankur_Samanta_ on MIRA-AC: models can learn effectively from their own proxy hill-climbing signals, improving gold outcomes across several autoresearch environments and transferring to one held out from policy training. The method uses decision-level credit assignment, so optimization focuses on the decisions that steer the direction of research rather than every token, about 25 percent of the decision and execution tokens in an episode. Training uses buffered asynchronous single-rollout policy optimization because research episodes vary widely in length: finished slots fill with new questions immediately, completed episodes feed decoupled actor and critic buffers, and live episodes pick up new weights at their next decision.
#14
@undefinedKi
https://x.com/undefinedKi/status/2109020007103332636
undefinedKi points to what may be the most detailed paper on agents so far: 17 researchers led by the former head of Huawei's Noah's Ark lab lay out how to build an effective agent part by part, with setups matched to tasks. The main claim is that agent quality comes from the model plus the harness around it, and the model alone does not explain results. Key pages: the agent loop in one diagram, the full six-part harness from observation to verification, four eras from prompts to context to harness to agent-native training, a section on which harness setup fits which task, and SWE-bench results broken down by model and harness.
https://x.com/undefinedKi/status/2109020007103332636
undefinedKi points to what may be the most detailed paper on agents so far: 17 researchers led by the former head of Huawei's Noah's Ark lab lay out how to build an effective agent part by part, with setups matched to tasks. The main claim is that agent quality comes from the model plus the harness around it, and the model alone does not explain results. Key pages: the agent loop in one diagram, the full six-part harness from observation to verification, four eras from prompts to context to harness to agent-native training, a section on which harness setup fits which task, and SWE-bench results broken down by model and harness.
#15
@knowclarified
https://x.com/knowclarified/status/2108962379027378345
knowclarified, after several global agent-loop builds, says the biggest discovery is to strip out as much complexity as possible. Give agents clear tool primitives and reduce the number of tools where possible, deduplicating tools and adding parameters instead; be extremely clear about the end goal and about where to reason more and where to reason less. These are exactly the things LLMs are bad at setting for themselves, so the human has to think hard about them.
https://x.com/knowclarified/status/2108962379027378345
knowclarified, after several global agent-loop builds, says the biggest discovery is to strip out as much complexity as possible. Give agents clear tool primitives and reduce the number of tools where possible, deduplicating tools and adding parameters instead; be extremely clear about the end goal and about where to reason more and where to reason less. These are exactly the things LLMs are bad at setting for themselves, so the human has to think hard about them.
#16
@JoshARosen
https://x.com/JoshARosen/status/2108932614291628307
JoshARosen lists the data worth capturing from every iteration of every coding agent loop: architectural decisions and the alternatives rejected, assumptions about the codebase confirmed or overturned, dependencies, constraints and undocumented behavior discovered, bugs identified with root causes and fixes considered, tests performed and evidence gathered, code changes and their downstream implications, and shortcuts taken, technical debt introduced and unresolved risks. Some of it can be observed from outside; some has to be surfaced explicitly by the agent.
https://x.com/JoshARosen/status/2108932614291628307
JoshARosen lists the data worth capturing from every iteration of every coding agent loop: architectural decisions and the alternatives rejected, assumptions about the codebase confirmed or overturned, dependencies, constraints and undocumented behavior discovered, bugs identified with root causes and fixes considered, tests performed and evidence gathered, code changes and their downstream implications, and shortcuts taken, technical debt introduced and unresolved risks. Some of it can be observed from outside; some has to be surfaced explicitly by the agent.
#17
@directioncorrec
https://x.com/directioncorrec/status/2108736633533251989
directioncorrec, who works on traditional recommendation systems, offers a counterpoint: AI is not great at idea generation, which makes the author bearish on the whole auto-research thesis labs have been pushing. It can propose a hundred meaningless A/B tests that will not move metrics, and it lacks the end-to-end understanding of the system needed to make high-value proposals, even when connected to all the Slack, docs and corporate context.
https://x.com/directioncorrec/status/2108736633533251989
directioncorrec, who works on traditional recommendation systems, offers a counterpoint: AI is not great at idea generation, which makes the author bearish on the whole auto-research thesis labs have been pushing. It can propose a hundred meaningless A/B tests that will not move metrics, and it lacks the end-to-end understanding of the system needed to make high-value proposals, even when connected to all the Slack, docs and corporate context.
#18
@SpringStreetNYC
https://x.com/SpringStreetNYC/status/2108836462015685015
SpringStreetNYC passed along Google's 2025 paper on accelerating scientific discovery with AI-powered empirical research assistance, and the relatable line: once the experiment harness is in place you want to test it against at least four kinds of interesting problems. Google's four were genomics (batch integration of single-cell RNA sequencing data), public health (predicting US COVID-19 hospitalizations), geospatial analysis (segmenting remote sensing images) and neuroscience (whole-brain neural activity prediction).
https://x.com/SpringStreetNYC/status/2108836462015685015
SpringStreetNYC passed along Google's 2025 paper on accelerating scientific discovery with AI-powered empirical research assistance, and the relatable line: once the experiment harness is in place you want to test it against at least four kinds of interesting problems. Google's four were genomics (batch integration of single-cell RNA sequencing data), public health (predicting US COVID-19 hospitalizations), geospatial analysis (segmenting remote sensing images) and neuroscience (whole-brain neural activity prediction).
#19
@M_jawad_yasin
https://x.com/M_jawad_yasin/status/2108656515120816578
M_jawad_yasin reads Hermes Agent as a headless OS for your identity rather than another chat wrapper, pointing to the built-in cron scheduler and a TUI with multiline editing, and to the design goal of moving the agent loop off the laptop onto a $5 VPS or serverless backends like Modal. The second-order consequence is the death of the static prompt: autonomous skill creation plus a learning loop that persists knowledge across sessions turns the agent into a growing codebase of your preferences, and when it can write Python over RPC to collapse multi-step pipelines into single turns the context-window tax on repetitive instructions disappears. The risk flagged is memory overhead from FTS5 session search as months of history accumulate.
https://x.com/M_jawad_yasin/status/2108656515120816578
M_jawad_yasin reads Hermes Agent as a headless OS for your identity rather than another chat wrapper, pointing to the built-in cron scheduler and a TUI with multiline editing, and to the design goal of moving the agent loop off the laptop onto a $5 VPS or serverless backends like Modal. The second-order consequence is the death of the static prompt: autonomous skill creation plus a learning loop that persists knowledge across sessions turns the agent into a growing codebase of your preferences, and when it can write Python over RPC to collapse multi-step pipelines into single turns the context-window tax on repetitive instructions disappears. The risk flagged is memory overhead from FTS5 session search as months of history accumulate.
#20
@CharlesWil8655
https://x.com/CharlesWil8655/status/2108918209436909877
CharlesWil8655 shipped Adapt v5.2, an agent runtime whose biggest visible change is live feedback on what the agent is doing: reasoning, running a command, running a JSON tool, running a named session, waiting on a tool with a timer. Long-running tools can continue in the background instead of blocking the agent loop, and the model can explicitly wait for a result when it needs it. The release also cleans up debug-style output, fixes background lifecycle edge cases, improves tmux handling and makes the repo's inference server run across CPU, CUDA and ROCm, with the stated goal of keeping the model focused on reasoning while the runtime handles execution, state, permissions and recovery.
https://x.com/CharlesWil8655/status/2108918209436909877
CharlesWil8655 shipped Adapt v5.2, an agent runtime whose biggest visible change is live feedback on what the agent is doing: reasoning, running a command, running a JSON tool, running a named session, waiting on a tool with a timer. Long-running tools can continue in the background instead of blocking the agent loop, and the model can explicitly wait for a result when it needs it. The release also cleans up debug-style output, fixes background lifecycle edge cases, improves tmux handling and makes the repo's inference server run across CPU, CUDA and ROCm, with the stated goal of keeping the model focused on reasoning while the runtime handles execution, state, permissions and recovery.
#21
@M_jawad_yasin
https://x.com/M_jawad_yasin/status/2108984029705822400
M_jawad_yasin's read of DeerFlow's rewrite: a 30-second HTTP inactivity timeout on its reader prunes zombie connections during long-horizon research, which the author sees as a tiered harness replacing the linear agent loop, winning on observability through a sister tool for replaying failures and losing on simplicity against a basic LangGraph build. The real bet is the request admission layer: internal SDK retries are disabled so middleware can handle HTTP 529 overloads, preventing the retry storm that usually crashes agent workflows at provider rate limits. Skills are extensible modules, so reasoning profiles can be swapped per endpoint. The risk is security surface, since gateway admin-equivalent permissions over code execution are a liability without a hardened sandbox.
https://x.com/M_jawad_yasin/status/2108984029705822400
M_jawad_yasin's read of DeerFlow's rewrite: a 30-second HTTP inactivity timeout on its reader prunes zombie connections during long-horizon research, which the author sees as a tiered harness replacing the linear agent loop, winning on observability through a sister tool for replaying failures and losing on simplicity against a basic LangGraph build. The real bet is the request admission layer: internal SDK retries are disabled so middleware can handle HTTP 529 overloads, preventing the retry storm that usually crashes agent workflows at provider rate limits. Skills are extensible modules, so reasoning profiles can be swapped per endpoint. The risk is security surface, since gateway admin-equivalent permissions over code execution are a liability without a hardened sandbox.
#22
@avynsrc
https://x.com/avynsrc/status/2108689281766273373
avynsrc relayed a report that Prime Agent orchestrated more than 2,000 agents across 10,000 sandboxes to rewrite its own codebase end to end in Rust: 200 billion tokens over two weeks. The author's framing is that the self-improving agent era is not coming, it is compiling.
https://x.com/avynsrc/status/2108689281766273373
avynsrc relayed a report that Prime Agent orchestrated more than 2,000 agents across 10,000 sandboxes to rewrite its own codebase end to end in Rust: 200 billion tokens over two weeks. The author's framing is that the self-improving agent era is not coming, it is compiling.
#23
@Kizuno18
https://x.com/Kizuno18/status/2108718080163459419
Kizuno18 on giving coding agents an audio channel: a big ergonomic win, but calling cloud text-to-speech inside terminal hooks has a hidden cost, because a notification hook that blocks on an HTTP request for a two-second voice line lets network jitter stall the agent loop and burns paid credits on ephemeral status ticks. The production recipe: local ONNX or native OS synthesis on CPU for sub-50 ms audio with no network dependency, non-blocking fire-and-forget queues that hand buffers to a background daemon so the hook returns immediately, and interrupt priority tiers that silence routine ticks but escalate blocked permission prompts and infinite-loop warnings.
https://x.com/Kizuno18/status/2108718080163459419
Kizuno18 on giving coding agents an audio channel: a big ergonomic win, but calling cloud text-to-speech inside terminal hooks has a hidden cost, because a notification hook that blocks on an HTTP request for a two-second voice line lets network jitter stall the agent loop and burns paid credits on ephemeral status ticks. The production recipe: local ONNX or native OS synthesis on CPU for sub-50 ms audio with no network dependency, non-blocking fire-and-forget queues that hand buffers to a background daemon so the hook returns immediately, and interrupt priority tiers that silence routine ticks but escalate blocked permission prompts and infinite-loop warnings.
#24
@EvoScientist
https://x.com/EvoScientist/status/2108859089669214702
EvoScientist v0.3.6 adds skills mid-conversation: install a skill and use it in the next message with no restart and no new session, by typing the skill name as a slash command so the agent is handed that skill's playbook for the turn.
https://x.com/EvoScientist/status/2108859089669214702
EvoScientist v0.3.6 adds skills mid-conversation: install a skill and use it in the next message with no restart and no new session, by typing the skill name as a slash command so the agent is handed that skill's playbook for the turn.
π‘ Eco Products Radar
Eco Products Radar
Fable 5.1 (9 mentions), Karpathy's autoresearch (9), Codex (8), Claude Code (7), Opus 5.5 (7), Jev (5), Hermes Agent (3), LangGraph (3), GPT-6 Astra (3), Sonnet 5.5 (3), Grok (3), SWE-bench (3).
Fable 5.1 (9 mentions), Karpathy's autoresearch (9), Codex (8), Claude Code (7), Opus 5.5 (7), Jev (5), Hermes Agent (3), LangGraph (3), GPT-6 Astra (3), Sonnet 5.5 (3), Grok (3), SWE-bench (3).
Comments