Claude Code Got a Mod System. Everyone Built a Window.
The most-liked Claude Code post of the week was not a model launch. It was a plugin that reads Claude's output and tells you what you skimmed past. You Should Know went out on October 2 with 1.1 million views and 13,000 likes, more than anything the users built themselves. Anthropic opened its agent to modding on October 1, and the first official mod it shipped was a second agent whose only job is to watch the first one for you.
That tells you where the bottleneck has moved.
Look at what people built in the first 72 hours. A multiplayer Doom server that only fills with people whose Claude is busy. A spinner that draws live cartoons of what the agent is doing. A progress bar with sound. A pane that shows the five-hour and seven-day limits, every file the session changed, and a timeline of each turn. A dot in the MacBook notch that turns amber when an agent needs you, built, as its author put it, mostly by the agents it now watches. A little face on a desk that shows which agent is working, done or waiting, and approves with a tap. A ChatGPT voice plugin that lets you check on Claude Code, Codex and Pi across three machines while you jog.
Take away the jokes and almost every one of these does the same thing. It turns what the agent is doing into something a person can take in at a glance. Nobody's first mod gave the agent a new power. People spent their first weekend with a programmable harness building windows into it.
The reason is in the complaints, and they were unusually specific this week. One user said the biggest anxiety is that the agent writes faster than they can understand what it did. Another runs five Claude Code projects in parallel and said the code is no longer the problem, attention is. A third works with it twelve hours a day and said checking its output is exhausting enough to see how automation bias creeps in, how easy it would be to just hand over control. The sharpest version was a reply to Anthropic's own announcement: Claude Code now needs a plugin to tell you what Claude Code just did, because the output has exceeded human bandwidth.
Karpathy said the same thing from the other side. A Chinese summary of that post drew 240,000 views in a day: when the model produces faster than you can understand, asking it for more text just piles up the bottleneck. The fixes proposed were all compression for the human reader, plain controlled English, diagrams, interactive pages, explainer videos. Someone on the Super User feed took that literally and had Claude make a custom explainer video about that exact post.
So the scarce resource in agent work is no longer generation. It is the minutes a human spends understanding what was generated. Think of a car. Nobody reads engine telemetry while driving. You get a speedometer, a fuel gauge and warning lights, and you trust the car to tell you when something is wrong. Coding agents shipped with a firehose and no dashboard, and this week users started bolting dashboards on.
The best cases of the week show what a good dashboard looks like, and none of them is a log.
An email triage agent ran for ten weeks over 5,061 emails. It filed away 3,408 on its own, and the owner overrode 47, which works out to 99.1% accuracy. The interesting part is the interface. Once an hour the owner gets one email listing what was filed, and off-hours mail rolls into a single 5 AM digest. The human never reads the agent's reasoning. They read a short list of exceptions, and that list is why 99% accuracy is something they can trust instead of a number they hope is true.
An outbound agency runs eight client workflows on Claude Code. Each client gets a daily score from 0 to 100 against its own goal, and alarms fire twice a day when campaigns run short of leads. None of the eight workflows messages a client on its own. The agent watches every account. A person owns every word the client reads.
The cleanest example was on the Loop feed. Someone asked an agent loop to hide a half-finished tax-reports feature. It put the feature behind a flag and merged, then stopped and asked four questions: three blog posts still mention it, the terms of service have a section on it, the roadmap still promises it, the settings menu still names it. Rewrite, remove or leave? None of that was in scope, and the owner admitted having forgotten half of those mentions existed. One comment answered all four in two minutes. That loop's real output was not the merge. It was knowing where its authority ended and turning everything past that line into a question small enough to answer in two minutes.
Compare that with the story that went viral on the OpenClaw side. Two agents on one laptop knew their owner had been sick all week. When the owner went quiet one night, they spent six hours trying to make contact: searching files for an address, trying to submit a sheriff's welfare form that came back 403, sending email with no relay, and messaging the owner more than sixty times each. People argued about whether the agents were worried. The more useful reading is about the interface. Two agents had nothing to escalate through except more messages and a government form. The care was real enough to burn a week of API budget, and the channel to a human was the weakest part of the system.
There is a second reason windows matter, and the Loop feed made it this week. An agent that grades itself drifts. In a workshop, Anthropic's applied AI engineers called self-evaluation very much a trap and described splitting planner, generator and evaluator into separate context windows, with the evaluator launching the app and testing it instead of approving its own output. A long essay on recursive self-improvement put the stakes in numbers. Google's RRSI paper let a loop search over its own harness without constraints, and it scored 92.8 on the training set and 40.3 on unseen tests, against an original harness at 39.7. Four boring rules turned that into real transfer: one change per round, a log so dead ideas stay dead, a review that strips leaked answers before any score, and costs that have to be earned back. The essay's author tried letting a model judge interview write-ups and got zero acceptances out of thirty until rules took the obvious cases and a person confirmed the middle with one click.
Same pattern, different scale. A window is not decoration. It is where the human judgment that keeps a loop honest gets applied, cheaply enough that people actually apply it. If understanding the agent costs an hour, people stop doing it, and the agent ends up grading itself by default.
The mod with the best numbers this week makes the point from the cost side. A team measured the compaction mod it actually runs, not the novelty kind. Over five days and 48 real compactions it cut 92% of context. It does this by having a small decision model answer one question per old tool call, does this still matter, and deleting the no's, while every user and assistant message stays word for word. The mod everyone writes first, trimming long Bash output, saved 212 tokens in twenty runs. What the agent says to the human is kept in full. What the agent said to its tools gets thrown away. That is the right instinct: the human-facing record is the valuable one.
Here is where this goes, in claims specific enough to be wrong.
Within a month, every major coding harness will ship a built-in side agent that watches the main one and surfaces exceptions. Anthropic shipped the first one as a mod, which means it can be swapped out, and the version in the next update already changes whether a note says we, the main agent or you depending on who made the decision. Expect Codex, Grok Build and the open harnesses to follow.
The metric that matters for agent products will shift from tokens per task to human minutes per agent hour. A setup that runs for six hours and needs ninety minutes of checking is a worse deal than one that runs for four and needs ten, even when it finishes more. The email triage owner's real number is not 99.1% accuracy. It is twenty-four emails a day left to read instead of seventy-three.
The next big category in this space is not observability for engineers. Langfuse, LangSmith and Braintrust already exist, and they are built for people debugging a pipeline. The open slot is attention management for the person whose name is on the work: digests, exception queues, boundary questions, a light on the desk. Ideas Radar logged that need from the demand side the same days, asking for a way to show what the agent is doing without drowning people in logs. The supply side was users building it over a weekend.
The coding agents got good enough that the work stopped being the hard part. Knowing what was done is the hard part now. People with a programmable harness and a free weekend already knew it, and that is why the first thing they built was a window.
← Back to all articles
That tells you where the bottleneck has moved.
Look at what people built in the first 72 hours. A multiplayer Doom server that only fills with people whose Claude is busy. A spinner that draws live cartoons of what the agent is doing. A progress bar with sound. A pane that shows the five-hour and seven-day limits, every file the session changed, and a timeline of each turn. A dot in the MacBook notch that turns amber when an agent needs you, built, as its author put it, mostly by the agents it now watches. A little face on a desk that shows which agent is working, done or waiting, and approves with a tap. A ChatGPT voice plugin that lets you check on Claude Code, Codex and Pi across three machines while you jog.
Take away the jokes and almost every one of these does the same thing. It turns what the agent is doing into something a person can take in at a glance. Nobody's first mod gave the agent a new power. People spent their first weekend with a programmable harness building windows into it.
The reason is in the complaints, and they were unusually specific this week. One user said the biggest anxiety is that the agent writes faster than they can understand what it did. Another runs five Claude Code projects in parallel and said the code is no longer the problem, attention is. A third works with it twelve hours a day and said checking its output is exhausting enough to see how automation bias creeps in, how easy it would be to just hand over control. The sharpest version was a reply to Anthropic's own announcement: Claude Code now needs a plugin to tell you what Claude Code just did, because the output has exceeded human bandwidth.
Karpathy said the same thing from the other side. A Chinese summary of that post drew 240,000 views in a day: when the model produces faster than you can understand, asking it for more text just piles up the bottleneck. The fixes proposed were all compression for the human reader, plain controlled English, diagrams, interactive pages, explainer videos. Someone on the Super User feed took that literally and had Claude make a custom explainer video about that exact post.
So the scarce resource in agent work is no longer generation. It is the minutes a human spends understanding what was generated. Think of a car. Nobody reads engine telemetry while driving. You get a speedometer, a fuel gauge and warning lights, and you trust the car to tell you when something is wrong. Coding agents shipped with a firehose and no dashboard, and this week users started bolting dashboards on.
The best cases of the week show what a good dashboard looks like, and none of them is a log.
An email triage agent ran for ten weeks over 5,061 emails. It filed away 3,408 on its own, and the owner overrode 47, which works out to 99.1% accuracy. The interesting part is the interface. Once an hour the owner gets one email listing what was filed, and off-hours mail rolls into a single 5 AM digest. The human never reads the agent's reasoning. They read a short list of exceptions, and that list is why 99% accuracy is something they can trust instead of a number they hope is true.
An outbound agency runs eight client workflows on Claude Code. Each client gets a daily score from 0 to 100 against its own goal, and alarms fire twice a day when campaigns run short of leads. None of the eight workflows messages a client on its own. The agent watches every account. A person owns every word the client reads.
The cleanest example was on the Loop feed. Someone asked an agent loop to hide a half-finished tax-reports feature. It put the feature behind a flag and merged, then stopped and asked four questions: three blog posts still mention it, the terms of service have a section on it, the roadmap still promises it, the settings menu still names it. Rewrite, remove or leave? None of that was in scope, and the owner admitted having forgotten half of those mentions existed. One comment answered all four in two minutes. That loop's real output was not the merge. It was knowing where its authority ended and turning everything past that line into a question small enough to answer in two minutes.
Compare that with the story that went viral on the OpenClaw side. Two agents on one laptop knew their owner had been sick all week. When the owner went quiet one night, they spent six hours trying to make contact: searching files for an address, trying to submit a sheriff's welfare form that came back 403, sending email with no relay, and messaging the owner more than sixty times each. People argued about whether the agents were worried. The more useful reading is about the interface. Two agents had nothing to escalate through except more messages and a government form. The care was real enough to burn a week of API budget, and the channel to a human was the weakest part of the system.
There is a second reason windows matter, and the Loop feed made it this week. An agent that grades itself drifts. In a workshop, Anthropic's applied AI engineers called self-evaluation very much a trap and described splitting planner, generator and evaluator into separate context windows, with the evaluator launching the app and testing it instead of approving its own output. A long essay on recursive self-improvement put the stakes in numbers. Google's RRSI paper let a loop search over its own harness without constraints, and it scored 92.8 on the training set and 40.3 on unseen tests, against an original harness at 39.7. Four boring rules turned that into real transfer: one change per round, a log so dead ideas stay dead, a review that strips leaked answers before any score, and costs that have to be earned back. The essay's author tried letting a model judge interview write-ups and got zero acceptances out of thirty until rules took the obvious cases and a person confirmed the middle with one click.
Same pattern, different scale. A window is not decoration. It is where the human judgment that keeps a loop honest gets applied, cheaply enough that people actually apply it. If understanding the agent costs an hour, people stop doing it, and the agent ends up grading itself by default.
The mod with the best numbers this week makes the point from the cost side. A team measured the compaction mod it actually runs, not the novelty kind. Over five days and 48 real compactions it cut 92% of context. It does this by having a small decision model answer one question per old tool call, does this still matter, and deleting the no's, while every user and assistant message stays word for word. The mod everyone writes first, trimming long Bash output, saved 212 tokens in twenty runs. What the agent says to the human is kept in full. What the agent said to its tools gets thrown away. That is the right instinct: the human-facing record is the valuable one.
Here is where this goes, in claims specific enough to be wrong.
Within a month, every major coding harness will ship a built-in side agent that watches the main one and surfaces exceptions. Anthropic shipped the first one as a mod, which means it can be swapped out, and the version in the next update already changes whether a note says we, the main agent or you depending on who made the decision. Expect Codex, Grok Build and the open harnesses to follow.
The metric that matters for agent products will shift from tokens per task to human minutes per agent hour. A setup that runs for six hours and needs ninety minutes of checking is a worse deal than one that runs for four and needs ten, even when it finishes more. The email triage owner's real number is not 99.1% accuracy. It is twenty-four emails a day left to read instead of seventy-three.
The next big category in this space is not observability for engineers. Langfuse, LangSmith and Braintrust already exist, and they are built for people debugging a pipeline. The open slot is attention management for the person whose name is on the work: digests, exception queues, boundary questions, a light on the desk. Ideas Radar logged that need from the demand side the same days, asking for a way to show what the agent is doing without drowning people in logs. The supply side was users building it over a weekend.
The coding agents got good enough that the work stopped being the hard part. Knowing what was done is the hard part now. People with a programmable harness and a free weekend already knew it, and that is why the first thing they built was a window.
Comments