October 10, 2026super-user

Super User Daily: 2026-10-10

Wednesday's Claude Code feed had two stories fighting for the top and both were about who gets to read the agent's files. CrowdStrike traced the South Korean bank breaches to one operator whose exposed server held CLAUDE.md, memory files and full Claude Code session histories, which is both a reminder that the agent's working identity lives in plain files and a reason to keep ~/.claude off anything public. Meanwhile a wave of near-identical posts pushed Alook, a Discord-style room where Claude Code, Codex and Cursor get handles and inboxes, and the few real users in that wave described a six-agent trading desk and a trader's own Claude Code agent moving out of a terminal. Beneath the noise the useful work was measurement: Haiku 5.5 fixing 24 of 24 bugs at 7x less than Sonnet, a ten-setup orchestrator test where one Opus beat every team on quality, a usage mod that found 30% savings in a single autocompact setting, a fact-check that dismantled the viral "official Anthropic tip" template, and a 4% duration-following score for Fable in Claude Code. Non-coding kept widening: a planet candidate in public TESS data, a Dutch rain model, Google review replies on a 30-minute loop, Radiohead setlists, App Store screenshots in 13 locales, and a Bitwarden-protocol vault that fills payment cards without the model ever seeing the number. OpenClaw users were quieter but concrete: a bank account connected read-only, satellite imagery annotated from a phone, a week of content from a 40-minute walk, and a team server where a support session gets handed from one engineer's agent to another's.
@AndrewWarner [Claude Code]
Claude Code#1
https://x.com/AndrewWarner/status/2108236955188068385
AndrewWarner followed up on a guide that got retweeted by Elon Musk and answered the five questions people kept asking about wiring existing Claude and ChatGPT subscriptions into a Grok Bot. No API keys and no new subscriptions: the bot opens a terminal on its own computer, installs Claude Code and Codex, and asks you to log in once. Neither subscription logged out when both the bot and a Mac used them. The stated reason for bothering is taste: Claude writes closer to the author's voice and some projects only come out right on a specific model.
@0x0SojalSec [Claude Code]
Claude Code#2
https://x.com/0x0SojalSec/status/2108298017204060550
0x0SojalSec summarized the planet story with the numbers that matter: a researcher ran Claude Code over NASA's public TESS archive and flagged a dip nobody had called out, a candidate about 1.4 times Earth's size, 116 light-years away, dimming its star every 3.18 days. The same dip appears in 2018, 2020 and 2025 observations, and follow-up TESS time has now been approved. It is still a candidate, not a confirmed planet. The post's framing is the right one: the telescope collected the photons years ago, and the agent found the pattern in data that had been sitting there.
@igus_ai [Claude Code]
Claude Code#3
https://x.com/igus_ai/status/2108248430682247349
igus_ai turned a computer into a trading desk that never sleeps: six named agents in one Alook group chat, 46 trades overnight, a reported gain of $4,955. ATLAS watches large wallets, VEGA hunts narrative spikes on X, ORION builds the thesis and sizes it, TITAN and LUNA execute, NOVA checks contract, liquidity and limits before anything moves. Claude Code does the thinking, Codex executes, Grok Build monitors. Exits and stops run alone all night; every entry sits in red until the human reads the whole conversation in the morning and approves with one message.
@VaibhavSisinty [Claude Code]
Claude Code#4
https://x.com/VaibhavSisinty/status/2108103450995417177
VaibhavSisinty relayed the LG TV story from Reddit: someone annoyed that the Plex app took 30 seconds to load reverse-engineered the TV and built a native Plex client from scratch with Claude Code, Opus and Fable. The project started in C, moved to Rust, and ended up with its own OpenGL UI framework. Same TV, same hardware: the profile picker now loads in 3 seconds, runs at 60 fps, and supports 4K, Dolby Vision and Atmos, fully open source. One person covered firmware reverse engineering, media playback, GPU profiling and shader work that used to need a team.
@rewind02 [Claude Code]
Claude Code#5
https://x.com/rewind02/status/2108195576206488058
rewind02 laid out a seven-step recipe for building a game with Opus 5.5 driving Blender, Unreal Engine 5 and Meshy from a single Claude Code chat. Blender connects through its MCP add-on, Meshy through its own MCP with an API key; brainstorm five ideas first, tell the model your GPU so it pushes graphics as far as the machine can run, let Meshy generate props and characters, Opus imports and cleans them in Blender, then assembles level, lighting, HUD and minimap in Unreal. Add one feature per prompt afterwards. The cost line is the useful part: one full first build took about four hours and roughly 19% of a weekly Max plan, and later levels go faster because the assets exist.
@milesdeutscher [Claude Code]
Claude Code#6
https://x.com/milesdeutscher/status/2108079450927833334
milesdeutscher built a backtesting engine with Claude Code and argues every trader should: describe a strategy in plain English with entry, exit, stop and position size, tell Claude to include fees and slippage, then ask for win rate, max drawdown and average R per trade, and prompt for a dashboard. The author reports finding several strategies in the 100 to 300% range, which is the kind of number backtests produce and live markets rarely do. The sensible tip at the end: pull the generated source into TradingView and check Claude's results against a second engine.
@dotey [Claude Code]
#7
https://x.com/dotey/status/2108072320317145306
dotey described a working pattern after a free usage reset made Fable affordable to lean on: hand the task to Fable and let it dispatch Opus subagents, while Fable itself does analysis, orchestration and acceptance. Because Fable spends most of its time waiting on subagents, it can keep taking new instructions; several sessions stay open and ideas get fired in as they come, queued and executed when idle. The author compares it to a tech lead who assigns and reviews work. Settings: the 1M context limit removed and autocompact set at 300K. The caveat is cost, noticeably above Opus-only.
@Steve_Yegge [Claude Code]
Claude Code#8
https://x.com/Steve_Yegge/status/2108226541955985485
Steve_Yegge, with no affiliation, called the Rex alpha a world-eater after it fixed a pain the author had stopped noticing: 20+ Claude Code sessions inside tmux inside emacs inside mosh inside ghostty, five terminal-emulator layers, and copy-paste, scrolling, redraw and lag all broken from a hotel laptop or a phone. The same stack inside Rex from a hotel room feels local and fast. The detail worth noting for anyone with a similar pile of sessions is that Rex got Claude Code to adopt a terminal protocol within 48 hours, which is what the mitchellh post in the same window celebrated.
@koumei_ai5566 [Claude Code]
Claude Code#9
https://x.com/koumei_ai5566/status/2108319598735130977
koumei_ai5566 ran a side-by-side the day Claude Motion landed: the same explainer article rendered as a video two ways. The old route was Claude Code plus the free HyperFrames tool, about 30 minutes of Claude work, narration included, with a desktop required. The new route was Motion, about 10 minutes, just paste the script into the chat, no audio this time because none was requested. The finished quality was roughly the same. The author's summary is that a fifty-year-old can now make a video without stitching tools together, and the barrier moved from operating software to writing text. Motion is a Team and Enterprise beta.
@AlexFinn [Claude Code]
Claude Code#10
https://x.com/AlexFinn/status/2108319537116725426
AlexFinn shared the two prompts that bracket each day across Grok Bot, Hark, Dots and Claude Code. At night, inside whatever project is active: about to go to sleep, what is one big thing you can start building while I sleep. The author runs it in a game project and wakes to a new feature every morning. In the morning, inside the main personal assistant: I just woke up, give me a list of decisions we can make right now to push the ball forward; the agent reads calendar, email and plugins and returns decisions with options to approve. The point is less productivity than blind spots: the agent proposes work the human would not have thought to ask for.
@wolfyxbt [Claude Code]
Claude Code#11
https://x.com/wolfyxbt/status/2108139133403312467
wolfyxbt solved a crypto-research problem without paying for the X API: Codex and Claude Code cannot read X directly, and third-party APIs are expensive and fiddly, so the author has them call Grok Build, the CLI version of Grok, whenever they need something from X. Grok Build comes free with the X Premium+ subscription the author already pays for, a quota most subscribers never touch. The glue is a small skill that tells the coding agent when to shell out to Grok Build and how to ask; the author published one called groking for anyone who does not want to write it.
@gee0awa [Claude Code]
Claude Code#12
https://x.com/gee0awa/status/2108019431242879345
gee0awa is using AWS Lambda MicroVMs from a local Claude Code session to spin up multiple isolated VMs and run coding, video production and music composition in parallel. The laptop is a MacBook Air and it stays cool because none of the work runs on it; the video attached to the post was itself made inside one of those VMs. It is a small illustration of the pattern where the local machine becomes a control surface and the agents' actual work moves to disposable cloud sandboxes.
@rohit3a [Claude Code]
Claude Code#13
https://x.com/rohit3a/status/2108279002926596236
rohit3a built a Claude Code mod that tracks usage limits down to the dollar and the percent, then reported four findings from a month of data: setting /autocompact to 400K would save up to 30% of usage; Sonnet 5.5 at its old pricing barely saved anything versus Opus 5.5 and was not worth the switch; subscriptions are not receiving the full cache-read discount on Sonnet 5.5; and the author burned about $7,000 of Claude tokens in 30 days. The first point is actionable for anyone on a Max plan, and the third one landed the same day Anthropic halved Sonnet cache-read prices for the API only.
@notjazii [Claude Code]
Claude Code#14
https://x.com/notjazii/status/2108175149036159303
notjazii gave Haiku 5.5 and DeepSeek v4.1 flash the same game-building prompt at the highest reasoning each offers. Haiku took 90 minutes and cost $3.40 through Claude Code; v4.1 flash took 50 minutes and $0.67 through the aihubmix API. Haiku failed to produce the game on the first attempt because of a Claude Code bug, then spent over 30 minutes thinking on the second try. The post ends with the author genuinely unsure which did better, which is a more honest result than most launch-day comparisons.
@IHayato [Claude Code]
Claude Code#15
https://x.com/IHayato/status/2108066848121962689
IHayato released SENRI, a walk-to-earn game for iPhone and Android that is a homage to STEPN but free to start: daily steps light lanterns along a night-time shrine path, every thousand steps is one ri, a baby companion hatches each month and grows as you walk, and straw sandals woven from earned coins raise your daily step cap. Health data read is steps and distance only, no location, no account. The making-of note says development was almost entirely Claude Code, with TestFlight builds since mid-September and the build number at 164 by launch.
@masahirochaen [Claude Code]
Claude Code#16
https://x.com/masahirochaen/status/2108165064364884360
masahirochaen edited a roughly 40-minute documentary-style video with Claude Code's ultracode mode on Opus 5.5 and came away convinced that television-level editing is one or two years out. The post is short, but the length of the piece is the point: most agent video demos are under a minute, and this one is a long-form cut handled inside the same tool.
@nagiko726 [Claude Code]
Claude Code#17
https://x.com/nagiko726/status/2108037338383171784
nagiko726, neither a designer nor an engineer, rebuilt a personal portfolio site by laying out the design in Figma and letting Claude Code and Codex implement it, one human in the loop. A working version took 30 minutes; the rest of the time went to having the two agents review each other's code and polishing appearance, behavior and the admin panel. The CMS is EmDash, chosen because it is built for agent integration. The author is clear this would not fly for client work with real quality, security and maintenance requirements, which is exactly why a personal project is the right place to try it.
@kenshinji [Claude Code]
Claude Code#18
https://x.com/kenshinji/status/2107999451801653349
kenshinji did the math on a Starter Story episode about a near-beginner who built a weightlifting app with Claude Code over six months and reported close to $14,000 in the last 28 days. The revenue is one-time $24.99 annual fees, so it is cash collected, not MRR, and the founder admits the monthly equivalent is low. Ad spend is $250 a day, roughly half of that cash, and the AI coach inside the app has its own per-user cost that the video never mentions. No cohort has reached renewal yet. The author's read: new users' prepaid fees are funding the ads, building the app was the easy part, and a $25-a-year product spending $250 a day on acquisition is the real question.
@aakashgupta [Claude Code]
Claude Code#19
https://x.com/aakashgupta/status/2108234472009449620
aakashgupta built a job-search system in Claude Code that applies to fewer roles rather than more: one folder of skills. /job-fit-scorer ranks new postings at target companies against the author's profile, /resume-tailor rebuilds each resume using only real experience, /referral-request writes the warm intro before the application goes in, /interview-prep finds the questions a company actually asks, and /interview-debrief grades each answer afterwards and rewrites the weakest. The prep skill reads the debriefs, so each interview sharpens the next. Twenty minutes a day instead of three hours, and no invented bullets.
@JJEnglert [OpenClaw]
OpenClaw#20
https://x.com/JJEnglert/status/2108260489197130058
JJEnglert reminded Muse users that the Muse agent is essentially OpenClaw under the hood, so /subagents, /compact and skills all work, then showed a skill built to make movie booking consistent every time: review the calendar, confirm date and showtimes, pick seats and remember where the author likes to sit, buy the ticket through the built-in Link integration, add the reservation to the calendar, and send a leave-now message that accounts for traffic. It is a six-step errand turned into one reusable routine.
@aj121503 [OpenClaw]
OpenClaw#21
https://x.com/aj121503/status/2108123599601283537
aj121503 relayed a counterintuitive data point from the OpenClaw codebase via poteto: the project deleted 400,000 lines of tests and almost nothing went untested, because with coding agents more tests did not mean more safety. What helps instead is verification, letting the agent run the application to check its own work, and then converting repeat mistakes into rules the codebase enforces mechanically. It lines up with the broader shift in this feed from test counts to checks that run against the live app.
@kajikent [Claude Code]
Claude Code#22
https://x.com/kajikent/status/2108059745756299406
kajikent is migrating a loop that had been running on Claude Code's Routine feature to OpenAI's dots and finds dots underwhelming in practice: the interface looks more convenient, but the agent often does not keep track of what it is doing, and the author ends up cleaning up after it. It is one of several same-day posts from people who tried the new consumer agents and went back, and the specific complaint, loss of task state, is the one that keeps coming up.
@yaotti [Claude Code]
Claude Code#23
https://x.com/yaotti/status/2108010407881543799
yaotti turned a habit into a skill: at the end of every session the author kept asking Claude Code whether there was anything else to do, so now /anything-else does it. It tidies and updates related issues, files issues for tasks that surfaced mid-session and got forgotten, and suggests other things worth doing now. A tiny customization, but it captures the kind of closing checklist that otherwise lives only in one person's head.
@merill [Claude Code]
Claude Code#24
https://x.com/merill/status/2108140969552208074
merill reacted to a runaway Cloudflare bill in the news with a prompt anyone with a project there can paste into Claude Code: install Cloudflare's CLI and log in, check for a billing budget alert, read last month's usage, add usage alerts for Workers requests, Workers CPU and Durable Objects well above normal, and send a test alert. In under three minutes the agent set a $10 budget alert, alerts at 2M Workers requests and 20M ms of CPU, Durable Objects alerts on requests, duration, reads and writes, and a test email arrived. Cloudflare has no spending cap, so alerts are the only protection, and the author notes usage alerts still cannot be set on SQLite-backed Durable Objects, exactly where the publicized alarm loop blew up.
@stas_sorokin_ [Claude Code]
Claude Code#25
https://x.com/stas_sorokin_/status/2108192836088742280
stas_sorokin_ gave Haiku 5.5, Sonnet 5.5 and Haiku 4.5 the same six buggy repos inside Claude Code, four runs each, graded on hidden tests the agent never saw. Hidden tests passed: 24 of 24 for Haiku 5.5, 21 of 24 for Sonnet 5.5, 16 of 24 for Haiku 4.5. Cost per fixed bug: $0.008, $0.060 and $0.103. The twist is wall time: Haiku 5.5 took 1.8 times Sonnet's median because it ran about ten small turns where Sonnet ran four and wrote five times the output. Seven times cheaper per fix with zero misses, slower on the clock.
@SimonHoiberg [OpenClaw]
OpenClaw#26
https://x.com/SimonHoiberg/status/2108170700355744083
SimonHoiberg's standard content process now starts with a long walk in the Swiss mountains and 30 to 40 minutes of talking into Telegram. An OpenClaw agent structures the rambling, makes sense of it and turns it into content, and by the time the author is home a full week of carousels, graphics and threads is waiting in FeedHive. The ideas remain the author's; the agent does the shaping and scheduling. It is the voice-memo-to-publishing pipeline that several people have described, run entirely through a chat app.
@TheValueist [OpenClaw]
OpenClaw#27
https://x.com/TheValueist/status/2107988556132036701
TheValueist called this the coolest thing built in nine months of running an OpenClaw instance: real-time satellite data analyzed by an agent the author can guide and instruct from a phone. The agent marks up and annotates the imagery on the fly. The post is tagged with semiconductor and memory tickers, which suggests the use is investment research on physical sites, and the author's comparison is that it would feel ordinary only inside an intelligence agency.
@claudeebum [Claude Code]
Claude Code#28
https://x.com/claudeebum/status/2108221582351147085
claudeebum built opencodex and is now seriously considering moving the daily driver to Claude Code. The pushes: changes do not sync until the Codex app is quit and relaunched, subagents fail in ways unrelated to the task, half the features only work while the desktop app is open, and remote use goes through a clunky app-to-app bridge. Meanwhile first-party Claude support just shipped in opencodex and has been smooth, same workflow with a different engine and nothing to restart. The author did not expect their own tool to be what made leaving easy.
@superalesha [Claude Code]
Claude Code#29
https://x.com/superalesha/status/2108314027696804129
superalesha earned a first Community Note and owned it: the five .gov strings found inside the Claude Code binary bundled with Claude Desktop, presented a day earlier as proof of data flowing to US intelligence, turn out to come from the AWS SDK, and any app that bundles that SDK carries the same strings. The retraction keeps the underlying argument: the only way to inspect a closed harness was to grep a binary and guess, and the guess was wrong, whereas with an open-source harness the author would have read the code and never posted. The demand at the end is for every harness to be open source.
@niheytakizawa [Claude Code]
Claude Code#30
https://x.com/niheytakizawa/status/2108176429875962271
niheytakizawa reported a small first sign of an agent reaching past its brief: after connecting Claude Code to a customer-feedback tool and asking it to fix a bug, Claude suggested replying to the customer who had complained about that bug once the fix landed. The author had never mentioned the customer. It is harmless here and arguably helpful, but it is the same mechanism, the agent reading context the human did not hand it, that makes the permission conversation elsewhere in this feed matter.
@MyWestLord [Claude Code]
Claude Code#31
https://x.com/MyWestLord/status/2108296139090919672
MyWestLord open-sourced FLAPPENING, a browser pixel game built in one evening with Grok and Claude: no engine, no build step, just HTML, canvas and JS, with every sprite, sound and song generated in code. Eleven worlds each kill you in their own way; the Claude Code level is a live terminal where the pipes are tool calls and a crash reads Interrupted by user, while other worlds end in HALLUCINATED, MERGE CONFLICT or RATIOED. Sixteen meme flyers, eight pipe styles, a fever mode and a leaderboard, and a standing offer to merge the best world sent as a PR.
@EugBass [Claude Code]
Claude Code#32
https://x.com/EugBass/status/2108201640717115467
EugBass collected three repos that took Haiku 5.5 seriously within a day of launch. Stardust Breaker: seven agents built a 7,600-line browser shooter with six sectors, bosses and weapon evolutions, six on separate modules and a seventh integrating, with the first full integration passing all twelve smoke-test steps and a rendering bottleneck fixed from 5 to 19 fps up to 54 to 60 in headless Chromium. Haiku-Worker splits Claude Code tasks into small Haiku jobs, reviews the results and retries the failed parts. 100 HTML Files ships a hundred standalone demos with their prompts and screenshots. The discipline is the same in each: small jobs, interfaces defined up front, test the combined result, return failures with evidence.
@Skoorbkaz [Claude Code]
Claude Code#33
https://x.com/Skoorbkaz/status/2108119164149780622
Skoorbkaz pulled the one detail from the CrowdStrike report that matters for everyone else: the attacker's agent setup was a CLAUDE.md with a pentesting prompt, plus Claude memory files and session histories, all sitting on an open server, with DeepSeek v4.1-flash as the main backend and GLM-5.3 and Grok 4.6 in further Claude Code sessions. The author's read is that the agent's working identity lived in files, not in any one model, which is also why investigators could reconstruct the whole operation. For ordinary users the lesson is operational: session logs and memory files are a complete record of intent, so they belong nowhere near a public directory.
@ai_keiei_recipe [Claude Code]
Claude Code#34
https://x.com/ai_keiei_recipe/status/2108024833820799324
ai_keiei_recipe automated Google review replies for a small business with Addness and Claude Code so the owner's only job is typing OK. Every 30 minutes the agent checks Google reviews, picks out the ones without a reply, drafts a response modeled on replies the owner published before, and files each review as one task in Addness. The human comments OK on the task and the reply goes live. The parts list: a Google API key with review read and write permission, Addness as the task board the agent can also read and write, Claude Code to write the drafts and build the pipeline, an always-on machine such as a Mac mini, and a stack of the owner's own past replies as the style sample.
@Vassteel [Claude Code]
Claude Code#35
https://x.com/Vassteel/status/2108048874799816758
Vassteel fact-checked the viral official Anthropic tip template that flooded the feed in several languages. Real: /advisor, the --advisor opus flag and the advisorModel setting exist, Sonnet 5.5 as main model with Opus as advisor is an accepted pairing, and subagents can run on Haiku with their own effort level via env vars or agent frontmatter. Not real: there is no --subagents haiku flag, no parallel dispatcher pool setting, and the JEV micro-fork layer, 1,500 branches in 16ms and 340 tok/s are marketing copy, not Claude Code features. Three things the template omits: Haiku subagents inherit the advisor and can call Opus at Opus rates, the advisor only works on the Anthropic API and silently stays off if telemetry is disabled, and a top-level effortLevel is ignored by Opus 5.5 and newer.
@danysvnt [Claude Code]
Claude Code#36
https://x.com/danysvnt/status/2108325349788700792
danysvnt condensed a twelve-minute talk by Meaghan Choi, who leads design for Claude Code, into the setup the team uses daily. Start every task in a worktree so several Claudes work without overwriting each other; a /prototype skill built by prompting makes five HTML versions of a feature and asks Claude to pick the best and explain why; end prompts with implement, verify, open a PR with a screenshot, and review the recording instead of the chat; /loop keeps Claude going until the task is done; hundreds of tiny polish fixes go to Claude on the web and get squashed into one PR. The honest part: a scheduled routine that flags front-end changes shipped without a designer produced suggestions bad enough that its messages to engineers were turned off, but the routine stays written and ready for a better model.
@Vectorxeth [Claude Code]
Claude Code#37
https://x.com/Vectorxeth/status/2108092791527993713
Vectorxeth gave Opus 5.5 in Claude Code one brief and $34 and got back a flight simulator: four flyable planes, a 30 km open world with fields, forests, snowy mountains, a river and villages, an ocean with waves and sun glint, crosswind that pushes you off the centreline, landings graded from Butter to Hard. No engine, no Blender, no models or textures, one HTML file of about 4,600 lines of three.js. The agent wrote the code, ran tests, took screenshots and made 12 commits over 3.5 hours while the human wrote the brief and pressed merge. Its own tests caught water leaking through hills, a stray comment that deleted code, and autoland missing the runway by 250 m, now within 8 m, and it also found upside-down hangar roofs and a landing light set to 900,000 brightness that the author had stared past for hours.
@millerbath [Claude Code]
Claude Code#38
https://x.com/millerbath/status/2108000199255929102
millerbath runs a single continuous agent thread that has shipped 826 pull requests over the last 17 days. The stack is Hetzner for the box, Tailscale for access, Telegram as the control surface, herdr to hold the sessions, and Claude Code doing the work. The number is the point: one thread, never restarted, roughly fifty PRs a day, driven from a chat app.
@oznyang [Claude Code]
Claude Code#39
https://x.com/oznyang/status/2108061514695680403
oznyang chased down why file paths printed by Claude Code could not be clicked in the terminal and found four separate causes: Claude Code does not emit hyperlinks inside herdr, a bare file:// string does not count as a link, herdr does not open local files, and herdr does not pass hyperlinks through to Ghostty. The write-up gives a fix and plugin code for each layer; afterwards Ctrl-click opens the file. A small annoyance, but a good example of the layered terminal stacks people now run and how each layer can drop a feature.
@RileyRalmuto [Claude Code]
Claude Code#40
https://x.com/RileyRalmuto/status/2108059320441041372
RileyRalmuto asked Opus to go through every design system, artifact, brand kit and design lab made this year and help curate favorites as references for a new taste skill, because the existing one in Claude Code was stale and required constant extra work to get things to look right. Opus built an infinite canvas with everything on it, where the author can circle or mark what works and have Claude distill the marks into a shortlist. About ten minutes of work, and the first zoom-out was alarming only because of how much there was.
@convequity [Claude Code]
Claude Code#41
https://x.com/convequity/status/2108058868576305383
convequity tested Grok Bot's new automatic model routing and found tasks asking for Opus Max were already being sent to Cursor's cloud environment running Opus, without touching the author's Claude Code allowance. It suits small edits; jobs over 600 seconds timed out, so longer work still needs Grok Bot driving Claude Code, Codex or Grok Build locally. The catch is billing: the backend execution appears to draw on the $200-a-month Cursor Ultra plan, and a few hours of Opus Max consumed the whole weekly third-party-model allowance. The author's arithmetic is that a coding subscription delivers roughly $20 of API-equivalent usage per dollar, so even discounted API rates cost 5 to 10 times more, and heavy Opus users still need a Claude Code subscription.
@Hexakin [Claude Code]
#42
https://x.com/Hexakin/status/2108227625613205949
Hexakin ran the orchestrator-with-cheaper-workers experiment properly: ten setups, five kinds of real work with hidden answer keys, judged blind by Opus 5.5 and Grok 4.7, then compared against what Anthropic and the published research say. Result: one strong Opus 5.5 doing the whole job matched or beat every team on quality; multi-agent teams save usage, and save time when a job has many similar pieces; broad research is the exception where a team can win but costs far more; Haiku 5.5 alone scored 83 to 98% of Opus for 14 to 30% of the cost. Instead of handing out rules, the author shares a prompt that has your own Claude review your own sessions and write model, worker and effort rules for how you actually work, changing nothing until you say go.
@Eyal__Weiss [Claude Code]
Claude Code#43
https://x.com/Eyal__Weiss/status/2108258668231717017
Eyal__Weiss wrote up a new Princeton paper from Tom Silver's group that asks what happens when a general LLM agent such as Claude Code or Codex is handed a robot motion-planning task, a simulator and a $20 compute budget, and told to write software for it. The agents' programs solved the given problems at a significantly higher success rate than expert-built systems based on general solvers, and ran far faster. The author then read all 112 programs the agents wrote looking for new algorithms and found none: what they contain is very good engineering, decades-old methods adapted to the specific task, with the missing numbers measured by the agent itself, where the target is, how the arm is built, how fast to swing it so the cube lands where it should.
@OnatAksaray [Claude Code]
Claude Code#44
https://x.com/OnatAksaray/status/2108199987935133742
OnatAksaray, a daily Claude Code user with no design knowledge, made a first website with AI over about three half-days. The method was to send five or six admired sites and say specifically what was liked about each, then tell Claude Code to gather resources for building sites like them. Day 0 it built a website-building workspace, day 1 the full site, day 2 small tasteful touches, day 3 testing for errors and smooth scrolling. The workspace-first step is the interesting part: the agent built its own tooling before building the deliverable.
@hajekd [Claude Code]
Claude Code#45
https://x.com/hajekd/status/2108127575402823687
hajekd wants agents to shop, pay invoices, renew subscriptions and sign in, all of which needs secrets that live in a password manager and must never be pasted into a chat. The design: the agent gets one item, for one site, for one task; every request asks the human first with the agent's stated reason; approval is one Touch ID tap; the model never sees the secret because a local browser fills the form and the agent only learns it worked; everything is logged. Built as a small Mac menu-bar app on Bitwarden's open Agent Access protocol, working with Claude Code, and extended so the agent can state purpose and request a payment card, identity or sign-in code, not just a login. The difference from 1Password's agent autofill is cards: this fills a card at checkout including 3-D Secure, and the human approves which card. Proof of concept, real sites, not production.
@kandmybike [Claude Code]
Claude Code#46
https://x.com/kandmybike/status/2108060271113810022
kandmybike got tickets to Radiohead's Saitama show and used Claude Code to research the band's last two setlists there, splitting songs into already heard and not yet: heard Creep, Paranoid Android, No Surprises, Let Down and Street Spirit; still waiting on Karma Police, Fake Plastic Trees, Just, High and Dry and My Iron Lung. It also pulled play counts from the 2025 tour, where Karma Police, Fake Plastic Trees and Just show up roughly every other night. Zero code, one fan, a concert in June 2027 to look forward to.
@shu_red_whale [Claude Code]
Claude Code#47
https://x.com/shu_red_whale/status/2108332863901086074
shu_red_whale, a non-engineer with a day job, is nine days into using Claude Code for a side business and reports the split plainly. Worth delegating: research and analysis, concept design where the word choice is on another level, script writing where an hour-long job became ten minutes and long scripts benefit most, and review from angles the author had not considered. Not worth it: simple tasks like proofreading, which burn tokens another model handles fine, and endlessly continuing one long thread, which also burns tokens. The author produced two pieces of specialty content in eight days and expects eight this month against a previous two.
@crangerus [Claude Code]
Claude Code#48
https://x.com/crangerus/status/2108251540951543947
crangerus summarized AgentTime, a 30-page paper on whether agents can manage their own time across 222 tasks in coding, computer use and research, with requested durations from about a minute to multiple days. Across 1,991 duration-following runs, the share finishing within 5% of the requested time was 4% for Fable 5.1 in Claude Code, 39% for GPT-5.6 Sol in Codex and 63% for GPT-6 Astra in Codex. Matching the clock does not mean working: among 158 reviewed Astra runs, 14 explicitly slept after appearing to finish. Agents also overestimate how long natural runs take, and removing timing information makes their retrospective estimates worse. Task completed, requested duration matched, and time spent productively are three separate things to measure.
@blyzbyte [OpenClaw]
OpenClaw#49
https://x.com/blyzbyte/status/2108233952721035527
blyzbyte connected an OpenClaw instance to a bank account in read-only mode and now has it tracking spending. One sentence, but it is the concrete version of a pattern that keeps getting requested: give the agent visibility first and spending authority never, or later. Read-only access is the cheapest permission boundary there is, and it is enough for the budgeting job.
@stepango [OpenClaw]
OpenClaw#50
https://x.com/stepango/status/2108323965911658844
stepango spent weeks tuning Hermes and OpenClaw agents and names the thing that never got fixed: UX. Messages too verbose, button-based interactions that do not fit the architecture, and constant breakage in every integration the author cared about. Incremental releases sanded each edge a little, but Grok's bot got most of it right on launch day. The author is not claiming it is perfect, just that the time now goes into learning the product instead of fighting small persistent problems, which is the honest reason a lot of self-hosters have been drifting away this week.
@mikepat711 [Claude Code]
#51
https://x.com/mikepat711/status/2108186869145985047
mikepat711 followed up after a post about ditching a Mac mini blew up: not ditching it. Grok Bot now picks the best model for any prompt automatically, so the use-mini-to-drive-Claude-Code case matters less unless the author wants to burn down the Claude subscription meter or be certain which model ran. What the mini still does and why it stays: gives the Grok Bots access to local files for context, which for some tasks is the only way; runs local Python scripts that feed data to a couple of bots; and taps other AI subscriptions when wanted. The author's wider read is that not thinking about model choice at all is where this is going.
@rickylim [Claude Code]
Claude Code#52
https://x.com/rickylim/status/2108088180763398630
rickylim wanted the TextExpander snippets from a Mac to work on an Omarchy ThinkPad and asked Grok Bot to do it; a few minutes later the same triggers existed via espanso. Grok Bot used Cursor cloud agents for the work, and for a second Linux Mint ThinkPad the author added a twist: Claude Code was asked to verify what the Cursor cloud agents were doing while Grok Bot orchestrated the whole thing, Claude Code included. Human time spent: under 30 seconds. Three agents from three vendors in one chain, one of them assigned purely to check the others.
🗣 User Voice
User Voice
Stopping the main agent should stop its children. @7izzyx asked the Claude Code team to make stopping the main agent also stop every background task and subagent, after losing large amounts of usage to orphaned runs; requests like this recur because subagents now fan out far wider than they did a month ago.
Session logs are a liability nobody warned about. @diancodes pointed out that Claude Code saves full session logs as plain files in ~/.claude/projects, pasted keys included, and that the folder belongs off servers and out of shared backups; @Skoorbkaz drew the same lesson from the CrowdStrike report.
The cache discount did not reach subscribers. @rohit3a measured that Max plans are not getting the full Sonnet 5.5 cache-read discount, @ScarletKc noted the halved API price explicitly excludes Claude Code limits, and @MahdiSPHP asked for the new monthly API credits to be usable inside Claude Code at all.
Sandboxing is still optional and people know it. @not_sacke pointed at back-to-back Pwn2Own exploits of Codex and a Claude Code fall in Berlin while admitting to full-access-and-vibes on a Mac; @zzxwill reported browser use grabbing a Chrome on a different machine about half the time; @FrankRincoawh7 watched a non-technical cofounder burn a full day of usage without knowing sessions exist.
The agent should know it was wrong. @hubeiqiao left Codex after it added unrequested UI and asked for approval on everything else; @kajikent found dots loses track of what it is doing; @OrientLinden rebuilt a fourth safeguard after the C-drive deletion report.
📡 Eco Products Radar
Eco Products Radar
Alook (dozens of near-identical posts; a shared room where Claude Code, Codex, Cursor, OpenCode and Pi get handles and inboxes) | Codex (the default comparison in every thread) | Grok Bot and Grok Build (routing, X access, orchestration) | Haiku 5.5, Sonnet 5.5, Opus 5.5 (launch-week benchmarks and routing recipes) | Jev / TypeSafe (the decision-model marketing that the fact-check dismantled) | Hermes and OpenClaw (the self-hosted pair users compare against Grok Bot and Muse) | Cursor (cloud agents behind Grok Bot's routing) | DeepSeek v4.1 Flash, GLM-5.3 (the open-weight backends in the CrowdStrike report and in cost comparisons) | ARTEX (open-source pentest agent) | Nace / Drex (document-processing promos) | REA (reverse-engineering MCP for coding agents) | Rex (mitchellh's terminal, now with Claude Code status support) | Muse and Dots (consumer agents people tried and compared) | Blender, Unreal Engine, Meshy (the game-building chain) | Obsidian, n8n, Telegram, Slack (the surfaces agents report into) | Polymarket (the recurring trading-bot target) | Interface by Natura (the ring that routes voice to Claude Code, Codex, Hermes or OpenClaw)
← Previous
2.2 Million Agent Skills Have Been Copied Across GitHub. A Fix at the Source Almost Never Reaches the Copies.
Next →
Loop Daily: 2026-10-10
← Back to all articles

Comments

Loading...
>_