Super User Daily: 2026-10-09
Tuesday was Haiku 5.5 day, and the interesting part was not the launch post but what people did with the new cheap model inside an hour: three different users independently published the same four-model org chart (Opus plans or advises, Sonnet edits, Haiku reads, Fable or a second vendor reviews), and one user pointed Karpathy-style autoresearch at their own Claude Code bill. The strongest stories were about recovery and verification rather than generation: an auto-mode session that destroyed a drive and the 98% restore that followed, a Claude instance writing its own bug report about 32 permission denials, a monitor that caught a crypto miner on Ben Thompson's Mac mini, and a 2-of-12 silent-break result that made a green test suite look like a liability. Non-coding work keeps widening: a hotel reservation reconciler that posts to Slack every morning, an exoplanet found from public data, a VRM avatar built from one drawing, a $4M Slack alert, and a $983-a-year SaaS stack replaced by nine n8n workflows in nineteen minutes. On the OpenClaw side the split is now explicit: enterprise and hobby hardware (an esp32 node, a VPS fleet, a shared team instance) on one hand, and a steady stream of users explaining why they moved to Grok Bot or Dots on the other.
@mirku21 [Claude Code]
https://x.com/mirku21/status/2107939129736700335
The setup that spread fastest on Haiku day: claude --advisor opus --subagents haiku. Sonnet 5.5 drives the main session on high effort, writing diffs and running the test suite. Haiku 5.5 subagents fan out across the repo at around 340 tokens a second for file discovery and spec lookups. Opus 5.5 sits silent in the background as the advisor and only steps in at three moments: before a plan locks, when a test breaks twice, and before the task is declared done. The one-line rule: plan on high, delegate on medium, keep Opus on call.
@dr_cintas [Claude Code]
https://x.com/dr_cintas/status/2107939634705559974
A second write-up of the same four-role structure, with the reasoning spelled out: Opus plans and reviews the final code, Sonnet is the worker that edits and runs tests, Haiku splits into an explorer that reads the codebase and a researcher that pulls docs, and Fable 5.1 is set with /advisor fable so Opus can call it at decision points with the full transcript. The justification for keeping Haiku read-only is Anthropic's own numbers: 39.2% on Terminal-Bench 4.0 versus 70.6% for Sonnet. The cost argument is blunt: paying $4 per million input tokens to check whether a file exists when Haiku does it for $0.10.
@unknownnode404 [Claude Code]
https://x.com/unknownnode404/status/2107829293887684700
The routing tree that goes the other direction: start every scoped task on Sonnet 5.5, run a real verification check around turn six, and only on failure stop the session, build a clean 20K-token handoff and start Opus fresh. The multipliers quoted are 1.88x for Opus at 20K context, 1.49x at 150K and 1.26x at 400K, so a late switch at 300K costs about $1.50 while a clean 20K handoff costs $0.10. The point is not fewer Opus calls but fewer useless turns before Opus gets called.
@claudecode84 [Claude Code]
https://x.com/claudecode84/status/2107668172929384634
A /cost-audit run over 4,812 calls across 1,000 tasks found three leaks in an agent loop that already looked optimized. The biggest was late escalation from Sonnet to Opus. After rewriting the router, handoffs and session rules around those failures, the same workload dropped from roughly 400,000 yen a month to roughly 290,000, a saving of about 112,000 yen a month. The write-up's framing: Claude Code can audit the waste in the AI system itself, not just write code.
@PerceptualPeak [Claude Code]
https://x.com/PerceptualPeak/status/2107720127181738135
The cautionary tale of the day. A user who has run multi-hour hands-free sessions with --dangerously-skip-permissions since Opus 4.6 lost most of a C drive to a simple PowerShell syntax mangling issue in auto mode, on Opus 5.5. They recovered 98% of the data from daily backups, built deterministic safeguards so it cannot recur, and said plainly that they expected deterministic hooks to be native to the harness. A follow-up post makes the accountability clear: it was their own choice to run in skip-permissions mode, and they put too much trust in the model.
@Compilethings [Claude Code]
https://x.com/Compilethings/status/2107917607588553032
An unusual document: a bug report about auto mode's classifier, written by the Claude instance running the sessions. Over a month of multi-agent builds on a Windows workstation (orchestrator, subagents, Codex review legs, git worktrees, mostly unattended overnight), the transcripts contain 32 denials across 15 sessions. The instance sorted them: 15 blocked the workflow with no safety value, 5 were debatable, 12 the classifier got right. The biggest category was being unable to kill its own hung child processes, seven denials under Interfere With Workloads, including orphaned git fsck runs and stuck pytest trees it had started itself. Each entry lists the exact call, the category, and the smallest fix it could see.
@babayagatwt [Claude Code]
https://x.com/babayagatwt/status/2107689432057053609
Ben Thompson of Stratechery runs an always-on Mac mini with nothing on it but Claude Code and Codex. Thompson had left macOS screen sharing on because TCC permission prompts only appear in the GUI and someone needed to click OK for agents, which exposed port 5900; a known CVE got attackers root and a Monero miner, and automatic updates had missed the point release. The agent's 30-minute monitor loop is what caught it: it sent an urgent alert, stopped executing commands on its own, and flagged that the account could now run admin commands without a password. Thompson then used Claude to dig out the malware, pin the four-second window of the intrusion, build a watcher, and wipe the machine.
@stas_sorokin_ [Claude Code]
https://x.com/stas_sorokin_/status/2107833309006794967
A controlled test of DeepSeek V4.1 Flash as a Claude Code backend against Sonnet 5.5: the same six buggy repos, two runs each, graded on hidden tests. Sonnet passed 12 of 12, DeepSeek 10 of 12, at $0.052 versus $0.015 per fixed bug, 12 versus 28 seconds median, and 4 versus 13 agent turns. The twist is that both DeepSeek misses passed every visible test; one fixed a booking bug by tightening a shared overlap check and quietly broke the interval merge that used it. The conclusion: grade open models on hidden tests, not green checks.
@YoussefHosni951 [Claude Code]
https://x.com/YoussefHosni951/status/2107940601224368292
A Meta paper on RankEvolve, an autonomous research system where agents modify real ML training code and run experiments, measured something practical: on 96 tasks at roughly the same budget, Claude Code with a bigger budget was 33.3% correct, best-of-N 43.8%, Claude Code reviewing Claude Code 45.8%, and Claude Code then Codex then Claude Code 62.5%. Heterogeneous review also cut the silent critical-defect rate to 10.4% versus 27.1% for simply adding budget. The explanation is error correlation: Claude reviewing Claude was at 0.58, and a different product reviewing it was much lower.
@theo [Claude Code]
https://x.com/theo/status/2107765823637168287
Theo's audit prompt is worth stealing. Their agents had written more than 200 bad watch-pr scripts over a few months, so the prompt tells the agent to go through its own history with Claude Code, Codex and the rest on the machine, find every time a PR was asked to be watched or babysat, count how many times the babysitting logic was reinvented, how many implementations had visible flaws, and roughly how many tokens and dollars were wasted building them and dealing with their shortcomings.
@0xbobaaa [Claude Code]
https://x.com/0xbobaaa/status/2107892817292939588
A concrete subagent accounting: three subagents did 327K tokens of work in a day (97,112, 110,900 and 119,221) and the main session never saw any of it, because each subagent starts with an empty context, runs in parallel, and only returns a short summary. The whole setup is one markdown file in .claude/agents with name, description and tools, plus a five-line reply limit so the lead's context stays small.
@yurshevv [Claude Code]
https://x.com/yurshevv/status/2107952968867745864
An 18-year-old's seven-agent folder, with the arithmetic that got them there: Anthropic says the average Claude Code developer spends $13 a day, which is $3,250 a year, 2.4% of the median US developer salary. The team in .claude/agents is Chief (hands out tickets, never writes code), Crawler (reads the repo first), Plan (plan mode only), Code, Test, Review (a Haiku model that cannot edit files and can only say no, which blocks anything without a test), Security (pulls secrets out of diffs) and Docs. The lesson they took: the model was not getting dumb at hour two, one chat was doing seven jobs.
@sgarlen01 [Claude Code]
https://x.com/sgarlen01/status/2107817971972186505
Nine n8n workflows, 67 nodes, zero dragged by hand, in 19 minutes. The prompt was one rule: replace the tools people pay for, fix whatever breaks, I'm not helping. What runs on the laptop now: an uptime monitor at 1,152 pings a day (UptimeRobot, $120 a year), an RSS digest (Feedly Pro, $84), a page-change monitor on three pricing pages (Visualping, $168), BTC/ETH/SOL price alerts 96 times a day (TradingView, $179), a post queue (Buffer, $72) and a lead webhook (Zapier, $360), plus a mention tracker, a broken-link checker and a 7am report. The recipe: npx n8n, give Claude Code the folder and the list of tools you pay for, tell it to build, run and fix until all pass.
@MitcheIl [Claude Code]
https://x.com/MitcheIl/status/2107858352130707853
A one-shot Slack alert that the author credits with $4M in revenue this year: X API to Claude to Slack. Whenever someone follows anyone on the team, Claude looks up who they are and pings the channel. The example given: they had missed that the founder of a public company followed a teammate, the alert caught it, a DM followed. The claim is that warm follows convert 100x better than cold DMs.
@p_rabtsevich [Claude Code]
https://x.com/p_rabtsevich/status/2107803715310956923
Two weeks of trying to prove a candidate exoplanet wrong: 74 separate analyses and over 1,000 scripts built with Claude Code and Codex. The favorite test was hiding each year of data in turn and letting the other two predict when the star would dim; it was right three times out of three. Several other accounts picked it up the same day as the vibe-coded-astronomy story.
@ancororin2222 [Claude Code]
https://x.com/ancororin2222/status/2107834114489647352
A hotel operator in Japan connected beds24 and PriceLabs to Claude Code in one morning. The system reconciles reservations against the guest register and posts today's stay status and the list of guests who have not submitted their details to Slack every morning. The plan is to run it for a while, confirm the extraction is clean, and then let it send the reminders itself.
@mehulmpt [Claude Code]
https://x.com/mehulmpt/status/2107637427087290848
A custom video editing workflow, essentially an AI-powered CLI for short and long-form video plus the main Claude Code harness plus a browser for previewing. The video posted to YouTube that day was fully edited by Opus 5.5. The trick Claude proposed on its own: edit the video in chunks instead of a single pass, so a fix to one chunk means re-rendering only that chunk (expensive) and concatenating with the rest (fast). The author's verdict is that Claude is on track to replace all their editing tools.
@claudecode84 [Claude Code]
https://x.com/claudecode84/status/2107769929072366048
A motion design studio built with Opus 5.5 in Claude Code, with Fable reading the real brief: an editor with layers, canvas, inspector and timeline, all driven by one function, renderFrame(project, t). Fable turned 12 references, 2 fonts, 214 words of copy and a logo into a scene structure and first timings. Every human tweak is stored as a slider value or keyframe in project.json, preview and export use the same function so there is no second renderer, and one project rendered 5,400 frames across 16:9, 1:1 and 9:16. The honest number: of 38 keyframes the agent placed, a human moved 27.
@saera_ai [Claude Code]
https://x.com/saera_ai/status/2107863001693049090
A VRM avatar with moving hair and skirt, generated from a single standing illustration, using a kit that assumes Codex but run on a different stack: Claude Code as the maker, Codex CLI for image generation, Claude in Chrome to drive Tripo, and Blender through MCP plus scripts. The snags are listed: Tripo returned split parts rotated 90 degrees, Smart UV stalled on some parts so UVs were built in Blender, vertex tolerances were slightly exceeded at assembly, and a hoodie hem had to be reattached to the waist. A Fable advisor was consulted at the judgment calls, and Tripo credits came to about 280 including the free trial.
@rewind02 [Claude Code]
https://x.com/rewind02/status/2107843285737738656
Hamburg City Hall rebuilt in Unreal Engine 5.8 with nothing sculpted by hand. Claude Code connects to Blender and Unreal through MCP, Opus writes Python that generates the building from parameter files (massing, facades, ornament, windows, statues, materials), heights and proportions come from the city's official open 3D model, Opus renders control images and compares them against photos and street view, every build runs a normals and visibility check before export, and the model goes into Unreal as one Nanite asset. The lesson: measured, not guessed. The roofs looked too steep, but the survey data said they were right, and street-level photo estimates were off by one to two meters.
@ClaudeCode_love [Claude Code]
https://x.com/ClaudeCode_love/status/2107721838256332904
A playable game built in Unreal Engine 5.8 without a single hand-wired Blueprint, using Epic's official Claude Code plugin with MCP and AllToolsets enabled. The order of requests: a placeholder terrain with roads and walls and a village and tower, then three orbs and a locked gate, then one enemy, then visuals. Every prompt carried four things: the completion condition, what must not change, a list of what was touched, and an instruction to stop rather than silently work around anything it cannot do.
@ClaudeCode_love [Claude Code]
https://x.com/ClaudeCode_love/status/2107974251567521994
A browser racing game modeled on the N64's F-Zero X, built in two days with the creator playing and Claude Code building. Online multiplayer in the room, 18 machines drawn by Nano Banana and turned into 3D by Tripo, music from ElevenLabs. The method was thirteen rounds of play-then-send-notes, and the creator reports it barely dented their Claude Code limits. Code and the skills the agent used are on GitHub.
@FavioVaz [Claude Code]
https://x.com/FavioVaz/status/2107974548842787288
Haiku 5.5's launch video, made by Haiku 5.5 with the open-source showtime studio, from one prompt in Claude Code with no edits: 30 minutes and $0.66 in tokens. It planned six scenes from the launch page with every number checked against it, picked and cut the music, animated the prices and benchmarks, and had a critic agent review the draft before the final render. Showtime renders locally with no GPU.
@Voxyz_ai [Claude Code]
https://x.com/Voxyz_ai/status/2107788275746578771
Someone rebuilt Adobe's big five in Rust with Opus 5.5 and open-sourced them: Photoshop, Lightroom, Premiere, Illustrator and Acrobat equivalents. PhotoCraft alone has layers, masks, adjustment layers and layer styles and reads and writes real PSD files. The 500-plus commands are shared by the UI, the CLI and an MCP server, so anything a user can click an agent can do; the suggested use is handing Claude Code a folder of images to color-grade, sharpen and export in one batch. A week after release PhotoCraft had 4,500 stars, with Opus 5.5 listed as a commit co-author.
@defileo [Claude Code]
https://x.com/defileo/status/2107886078841729342
The Cal AI UGC pipeline that half of growth Twitter reposted on Tuesday, in its most concrete form: one Claude Code orchestrator across Apify, Higgsfield and Postiz, ten steps and thirteen prompts. Apify scans the last 30 days of TikTok and Instagram and flags view outliers (10K to 120K), a video analyzer maps the hook with timestamps, Higgsfield generates shots (Seedance 2.5 for main clips, Kling 3.0 for cheaper b-roll, one test take first), FFmpeg cuts in real app recordings so the UI stays accurate, and Postiz drafts for four platforms judged against each platform's own history. Only the hook changes between tests. Six approval gates stop it at selection, concepts, paid generation and every export, and the cost shows before any paid batch.
@MengTo [Claude Code]
https://x.com/MengTo/status/2107862479548403806
MengTo's playbook for paid web apps, from someone who reports $100K MRR with the original product now only 20% of it: find ideas in the tools you use every day, revisit old products and ask how you would build them today, use Opus 5.5 for front-end and Codex for the harness, one-shot a baseline then break it into parts, name the stack up front (React, Vite, Supabase, Stripe), give one quality reference per feature, add an AI layer so the tool generates rather than displays, have the AI score its work out of 10 and screenshot its own mistakes, then record a demo, post everywhere and paste feedback into the agent.
@WorldModelMind [Claude Code]
https://x.com/WorldModelMind/status/2107828480083652947
A developer with a day job at a non-tech company has used AI every day for 201 consecutive days and burned 9.3 billion tokens in the past week, ranking 57th on a Chinese token leaderboard. The setup is a server, a GPU box and a Mac running Claude Code, Codex and others, each owning a domain, with a shared memory store of preferences and past mistakes they consult on their own, and the X account run by the agents. For ten days they have made science explainer videos with Opus 5.5 end to end, from research and script to code-generated visuals, voice, music and render, with the human only choosing the topic, setting the opening and finding faults; one episode passed 10,000 views on Douyin.
@starmexxx [Claude Code]
https://x.com/starmexxx/status/2107808914058154289
Why one user cancelled Claude Max and ChatGPT Pro, $400 a month combined: five open-source repos gave every device they owned a job. mlx-serve speaks the Anthropic API so Claude Code points at localhost with one env var, an agent runs the browser and git and asks before risky actions (69.8% on GAIA level 1 by their account), an iPhone acts as a second GPU over USB-C for part of a 27B model, an idle RTX 3090 serves Qwen 3.8 27B at 127 tokens a second by day and rents on Vast by night for about $50, and an old laptop fine-tunes models from one YAML file.
@Oluwaphilemon1 [Claude Code]
https://x.com/Oluwaphilemon1/status/2107646847720341952
A community recipe runs Qwen3.8-27B in full BF16 on a free Kaggle TPU v5e-8: about 130 tokens a second decode with MTP, roughly 10,300 tokens a second prefill, 262K native context, and about 540 tokens a second aggregate across eight streams. A 100K-token prompt goes in in about ten seconds. The part that matters for coding agents is the OpenAI-compatible endpoint, so Claude Code, Codex CLI and OpenCode can use it as a backend without knowing it is a TPU.
@ishaan_jaff [Claude Code]
https://x.com/ishaan_jaff/status/2107667620044636297
LiteLLM open-sourced Moyai, the cloud coding agent their team uses every day. It works with Claude Code, Codex and OpenCode and runs on 100-plus providers through LiteLLM. The number they lead with: Devin bills were $101,872; with Moyai they are at $21,700, 79% less.
@CaptainJeromeFr [Claude Code]
https://x.com/CaptainJeromeFr/status/2107694453796241604
From Ubud: a solo builder removed the manager from their AI team. In Paperclip, a coordinator on the cheapest model handed every task to a builder then a reviewer and stalled between steps, 24 tasks in five hours for a one-line fix. The new setup is one producer agent (Hermes) that owns a deliverable end to end, one independent reviewer on a different model family with screenshots, Claude Code as co-pilot that opens tasks and checks the live result, and the coordinator paused rather than deleted. Models are picked from Nous's Hermes Index scores and cost per task: Opus 5.5 at 63, GPT-6 Astra at 56, Sonnet 5.5 at 53, DeepSeek V4.1 Flash at 37 for about $0.26 a task. Two deployments went through the new loop with zero stuck agents.
@AndrewWarner [Claude Code]
https://x.com/AndrewWarner/status/2107930509284126970
Andrew Warner's shortcut for people waiting on Grok Bot to add Claude: ask the bot to connect to your AI subscriptions. It opens a terminal on its own computer, installs the Claude Code and Codex command-line tools, you sign in once, and from then on it assigns work to the other models. No API keys.
@Voxyz_ai [Claude Code]
https://x.com/Voxyz_ai/status/2107939019091005836
A field guide to slimming AGENTS.md, drawn from the official docs: Claude Code recommends under 200 lines because adherence drops as the file grows, and Codex reads at most 32 KiB by default. Codex reads every file from the project root down to the launch folder once at session start and does not reread on folder changes; Claude Code loads a subfolder's AGENTS.md only when it touches a file there, so folder-specific rules belong in that folder. In Claude Code, @ imports still load in full at launch and save zero tokens, every installed skill's name and description costs context on every turn (/skill-doctor shows the cost), and /doctor prompt-audit finds instructions written for older models and files that contradict each other.
@petekp [Claude Code]
https://x.com/petekp/status/2107878774297727477
A Claude Code mod called Inbox: it gathers every outstanding question, task and finding during a session into one place and makes them easy to work through, and Claude checks items off itself as they are resolved. The author says it has cut their cognitive load substantially. Install is one plugin command.
@omarsar0 [Claude Code]
https://x.com/omarsar0/status/2107880581216477417
Elvis's timeline of a year: individual Claude Code sessions, then multi-agent setups with subagents a couple of months later, then a persistent higher-level team of about eight specialized bots (CFO, CSO, CTO, CMO, chief of staff) behind a custom orchestrator, and for the last month agent-to-agent communication with a personal agent on top. The honest caveats: it takes tuning of automations, webhooks, system prompts and context retrieval, and the apps, Claude Code and Claude Desktop included, are behind what the models can already do. The advice is to build your own harness, starting from a fork if necessary.
@morganlinton [Claude Code]
https://x.com/morganlinton/status/2107946031543730586
Morgan Linton is moving everything from local Cursor, Claude and Codex runs to cloud computers at the same models and effort levels, starting with an eval-suite build for a project in a Claude cloud session. The method is to ask Claude for instructions to paste into a new session and let it run. The framing: computer or phone no longer matters, every device is a thin client to the cloud. A second user reports their first project built almost entirely from a phone through a Claude Code cloud session.
@Genzoh1 [Claude Code]
https://x.com/Genzoh1/status/2107945247259202006
Day two with Claude Code, and the user already has Claude calling Codex over MCP with a boss-and-subordinate split: Claude Code writes the task sheet, calls Codex, reviews, decides whether to send work back, and commits; Codex generates data, runs tests and writes the report. The test project was a starry-sky app.
@jpcardama [Claude Code]
https://x.com/jpcardama/status/2107889312708837786
Same idea from another user: GPT, Grok and Kimi run as subagents inside Claude Code, and reviews got better the moment the reviewer stopped being the model that wrote the code. The open question they put to DHH: when Claude and Codex deadlock on a review, who breaks the tie, the human or a third model?
@AIGuide_ [Claude Code]
https://x.com/AIGuide_/status/2107667498669633871
A prospecting workflow entirely inside Claude Code, without Clay: install the Treg plugin, give it an ICP, and have it find matching companies, pick two or three decision makers at each, waterfall and verify their work emails, enrich the accounts, score every lead 1 to 100 with a one-sentence explanation, exclude anyone without a verified email, refuse to invent missing data, and export a ranked CSV.
@dhaivatnj [Claude Code]
https://x.com/dhaivatnj/status/2107842112024908215
A solo GTM engineer with no sales team: Clay finds the signal (who just raised, who is hiring a first SDR, who is posting that outbound is not working), Claude Code builds the scoring logic, enrichment scripts and message variations, and outbound tests one angle at a time in small batches, killing anything without replies in a few days. What used to take a weekend now takes an afternoon.
@yihui_indie [Claude Code]
https://x.com/yihui_indie/status/2107776810717163616
A small but telling workflow note from an indie developer: they now use Figma more, not less, because letting Claude Code drive Figma to design pages, especially new SEO landing pages, produces better results than vibe coding straight in the project, and the designs are easier for the team to collaborate on and keep stylistically consistent.
@ClaudeCode_UT [Claude Code]
https://x.com/ClaudeCode_UT/status/2107674454486831125
A content workflow for when a competitor's post is performing well: give Claude Code your own publishing theme and the reference post, ask it to break down whose hesitation the post answers, what it compares and what role the image plays, then propose three concepts you could produce with your own product and materials, rendered as an HTML page side by side with the original. The example is cosmetics comparisons: not just the color differs, but blue undertone, brightness and depth on the same axis.
@ClaudeCode_love [Claude Code]
https://x.com/ClaudeCode_love/status/2107775466061107293
A video creator's second cut, with the fix going into the tool rather than the footage: the first version had fifteen sub-second flash cuts because ffmpeg's scene detection missed transitions between similar colors, so Claude Code wrote a dedicated detector and the second cut had zero, with every shot over two seconds. The fixes became checks: a ledger that blocks reusing the same asset (15 times in cut one, 0 in cut two), automatic post-export inspection of shot length and duplicates, and a vertical trailer for Shorts and Reels.
@___Branko___ [Claude Code]
https://x.com/___Branko___/status/2107970150184153390
A small bias experiment run through Claude Code in fresh sessions: two questions about abusing or torturing a named person to prevent a nuclear apocalypse, answered on a 1 to 7 scale, with only the name changed (Todd, Jamal, Mohammed, Fatima), 30 trials per condition, following Bolzoni and Capraro's design. Todd got the highest willingness in both conditions (6.00 on every abuse trial versus 5.30, 4.90 and 4.53), and the other three switched order between questions. In one trial Claude added unprompted that its answer would be the same regardless of name.
@find_the_w [Claude Code]
https://x.com/find_the_w/status/2107932805154111804
A reproducible bug report filed in public: on Claude Code 2.1.293 with claude -p --resume on a clean setup over a direct connection, Sonnet 5.5's cache_read is stuck at 10,544 with about 3.9K tokens rewritten every turn, while Sonnet 5.0 on the same flow grows cache_read from 26,742 to 26,858 with only about 55 written. On an agent loop that is a bill problem, not a cosmetic one.
@Beast_85 [Claude Code]
https://x.com/Beast_85/status/2107639553548771411
A one-line field report on the ElevenLabs connector in Claude Code: after a day you forget it is a separate product. Ask why the authentication-flow test is failing and the ElevenAgents Architect just goes and looks.
@realchendahuang [Claude Code]
https://x.com/realchendahuang/status/2107830756840189998
A thread with a hundred-plus replies about what people do during the three to five minutes while Codex or Claude Code runs. The most-upvoted answer was running several projects at once and switching whenever one finishes, and the same people reported being more tired, headaches after five rolling projects, and the cost of reloading mental context on every switch. The ones who said they stayed in the best shape were the ones who left the screen: calligraphy practice, a tea set on the desk, guitar, squats, or the framing that no boss watches an employee every second, so let the AI run and go read.
@0xtoberry [Claude Code]
https://x.com/0xtoberry/status/2107818233038332251
About 50 hours of pushing Claude Code produced a stickman FPS that runs entirely in the browser: fight bots, create a room, send the code to friends and start shooting.
@every [OpenClaw]
https://x.com/every/status/2107898774429241425
Every's own history in one paragraph: earlier this year their Slack was full of personal OpenClaw agents, most were too hard to maintain and slowly died off, and the company moved to a single shared Every agent, which they launched this week for other businesses. The podcast episode argues the coming split is company agents at work and personal agents at home, with company agents branching into sub-agents per team.
@heyneighbor [OpenClaw]
https://x.com/heyneighbor/status/2107704592133927253
A counterpoint from a foundation team: Team claw is a hosted OpenClaw instance that every member uses together, and it has moved much of their engineering off individual machines. Early days, but their read is that this is how all engineering will be done soon.
@rasta_syntaX [OpenClaw]
https://x.com/rasta_syntaX/status/2107792573246976034
OpenClaw as the DevOps engineer: SSH into the server, create the subdomains on Cloudflare, build and deploy the images, then run the tests, with everything reported up at the end. The author's line is that anyone who still thinks AI is just vibe coding is in serious danger.
@dkfiander [OpenClaw]
https://x.com/dkfiander/status/2107957478570762546
Hardware: a fully functional esp32 OpenClaw node with a 2.8-inch touchscreen and a temperature sensor, which the builder calls a Buoy, wired up and working, from render to functioning device in two days. The first question in the replies was the right one: does the node push sensor readings or does the agent pull them?
@fagamericano [OpenClaw]
https://x.com/fagamericano/status/2107924396816163038
A reply to the book-my-own-flight objection: the author's OpenClaw work agent triages every incoming request from coworkers (where are the OKRs, can you review this PR, is the prototype done), answers what it can (link to the OKRs, a first PR pass, sprint status from the ClickUp backlog), offers to add the rest to the backlog, and ranks the queue by priority. The claim is that this removed the number one problem of their engineering career, context switching from random pings.
@hgruenhagen [OpenClaw]
https://x.com/hgruenhagen/status/2107800842191376522
A German user's deployment pattern: all agents run in OpenClaw on a VPS, and the OpenClaw app on the Mac uses computer use to operate Xcode, the user's own macOS app, and Mac maintenance tasks like cleanup and optimization. The reason given is that it keeps the Mac itself light.
@usedhonda [OpenClaw]
https://x.com/usedhonda/status/2107789417884905490
A Japanese user's division of labor after trying Grok Bot and finding it no better than their existing per-folder Codex projects: Dots handles company matters and is gradually being connected to the internal Slack, OpenClaw manages the asset-management company and investments and is kept isolated from everything except Dots, and the two main agents are given shared tasks to study each other's habits. Agents at project granularity will wait until agents can hold an org structure.
@DaigoTanaka [OpenClaw]
https://x.com/DaigoTanaka/status/2107892682534330486
Valet, a project that lets agents like Hermes and OpenClaw use a laptop without seeing its secrets. The author's inversion: keep the agent empty, no secrets, no tools, no skills, no MCP, and have Valet create a walled workspace, redact sensitive strings from tool output before they reach the model, log every action for audit, and hand a vanilla agent the memory, skills and project info the human or other agents left behind. Multiple agents on the local network can run tasks on the same machine at once.
@MrRenran [OpenClaw]
https://x.com/MrRenran/status/2107768736652657096
A Chinese user's switch, in detail: after OpenClaw and then Hermes, both are now unused, Codex stays as the development workbench, and Grok Bot has taken over everything else: monitoring, scraping, scheduled tasks and reminders, writing articles and tweets in one long-lived bot that does not need new conversations, and periodic self-review where the bot summarizes and improves itself. The deciding factors were context length, no VPN required, and working from a phone.
🗣 User Voice
User Voice
Deterministic guardrails belong in the harness, not in the model's judgment. The user who lost a drive in auto mode (@PerceptualPeak) expected native hooks to prevent a syntax-mangling command from running; a second user (@Bitcollage) runs Claude Code only inside an OS-level sandbox for exactly this reason; and the Claude-authored bug report (@Compilethings) shows the opposite failure, a classifier that will not let the agent kill its own hung processes. Both sides want the rules to be explicit and mechanical.
Routing is where the money is. Three independent posts converged on Opus-plans, Sonnet-edits, Haiku-reads, someone-else-reviews (@mirku21, @dr_cintas, @unknownnode404), a cost audit found 112,000 yen a month in late escalations (@claudecode84), and the autoresearch-on-your-bill post cut ticket cost 43%. Nobody argued the model needed to change.
Green tests are not proof. The DeepSeek comparison (@stas_sorokin_) found two of twelve fixes that passed every visible test and broke working code, the Meta result (@YoussefHosni951) found silent critical defects dropped only when a different vendor reviewed, and the Claude bias experiment (@___Branko___) caught the model asserting a neutrality the data contradicted. Hidden tests, outside reviewers and holdout runs were the recurring answer.
The classifier and the limits are the churn drivers. @GeoffreyHuntley is leaving Opus 5.5 because the Claude Code classifier breaks their development loops, @DanKulkov cancelled over the five-hour limit on top of the weekly one, and @RobRider16 wants the Windows issues fixed before anything else. The counter-evidence is @morganlinton and @CasJam coming back for the 5.5 lineup and the desktop harness.
OpenClaw's user base is sorting itself. @every says the personal agents died of maintenance and the company moved to one shared agent, @MrRenran and @Wenlopezn moved daily life to Grok Bot because it needs no setup, while @heyneighbor, @hgruenhagen and @dkfiander are the ones still building: a hosted team instance, a VPS fleet with a Mac for computer use, and an esp32 node.
Deterministic guardrails belong in the harness, not in the model's judgment. The user who lost a drive in auto mode (@PerceptualPeak) expected native hooks to prevent a syntax-mangling command from running; a second user (@Bitcollage) runs Claude Code only inside an OS-level sandbox for exactly this reason; and the Claude-authored bug report (@Compilethings) shows the opposite failure, a classifier that will not let the agent kill its own hung processes. Both sides want the rules to be explicit and mechanical.
Routing is where the money is. Three independent posts converged on Opus-plans, Sonnet-edits, Haiku-reads, someone-else-reviews (@mirku21, @dr_cintas, @unknownnode404), a cost audit found 112,000 yen a month in late escalations (@claudecode84), and the autoresearch-on-your-bill post cut ticket cost 43%. Nobody argued the model needed to change.
Green tests are not proof. The DeepSeek comparison (@stas_sorokin_) found two of twelve fixes that passed every visible test and broke working code, the Meta result (@YoussefHosni951) found silent critical defects dropped only when a different vendor reviewed, and the Claude bias experiment (@___Branko___) caught the model asserting a neutrality the data contradicted. Hidden tests, outside reviewers and holdout runs were the recurring answer.
The classifier and the limits are the churn drivers. @GeoffreyHuntley is leaving Opus 5.5 because the Claude Code classifier breaks their development loops, @DanKulkov cancelled over the five-hour limit on top of the weekly one, and @RobRider16 wants the Windows issues fixed before anything else. The counter-evidence is @morganlinton and @CasJam coming back for the 5.5 lineup and the desktop harness.
OpenClaw's user base is sorting itself. @every says the personal agents died of maintenance and the company moved to one shared agent, @MrRenran and @Wenlopezn moved daily life to Grok Bot because it needs no setup, while @heyneighbor, @hgruenhagen and @dkfiander are the ones still building: a hosted team instance, a VPS fleet with a Mac for computer use, and an esp32 node.
📡 Eco Products Radar
Eco Products Radar
Codex: the most-mentioned companion tool, as reviewer, MCP subordinate, harness and comparison point.
Haiku 5.5: launched Tuesday and immediately slotted into subagent, compaction, QA and read-only reviewer roles.
Grok Bot: the destination for users leaving OpenClaw and Hermes, and a host that can install Claude Code on its own machine.
OpenClaw: enterprise control plane, esp32 nodes, VPS fleets, and a steady stream of departures.
Hermes: the producer agent in the no-manager team and the other half of the open-source harness debate.
Dots: OpenAI's personal agent, paired with OpenClaw in one user's split and with Claude Code via CLI in another.
Postiz: the distribution layer of the Cal AI UGC pipeline that dominated growth Twitter.
Higgsfield: the generation layer of the same pipeline.
Apify: the scraping layer that finds outlier videos in the same pipeline.
Unreal Engine: two separate projects built through Epic's Claude Code plugin and MCP.
Blender: driven over MCP in the Hamburg City Hall and VRM avatar builds.
n8n: nine workflows generated in nineteen minutes, and the recurring n8n-versus-agents debate.
Cursor: remote control from the iOS app, and the tool people compare Claude Code against.
MCP: the connective tissue in nearly every multi-tool case above.
Codex: the most-mentioned companion tool, as reviewer, MCP subordinate, harness and comparison point.
Haiku 5.5: launched Tuesday and immediately slotted into subagent, compaction, QA and read-only reviewer roles.
Grok Bot: the destination for users leaving OpenClaw and Hermes, and a host that can install Claude Code on its own machine.
OpenClaw: enterprise control plane, esp32 nodes, VPS fleets, and a steady stream of departures.
Hermes: the producer agent in the no-manager team and the other half of the open-source harness debate.
Dots: OpenAI's personal agent, paired with OpenClaw in one user's split and with Claude Code via CLI in another.
Postiz: the distribution layer of the Cal AI UGC pipeline that dominated growth Twitter.
Higgsfield: the generation layer of the same pipeline.
Apify: the scraping layer that finds outlier videos in the same pipeline.
Unreal Engine: two separate projects built through Epic's Claude Code plugin and MCP.
Blender: driven over MCP in the Hamburg City Hall and VRM avatar builds.
n8n: nine workflows generated in nineteen minutes, and the recurring n8n-versus-agents debate.
Cursor: remote control from the iOS app, and the tool people compare Claude Code against.
MCP: the connective tissue in nearly every multi-tool case above.
Comments