Latest · 最新
Sep 18, 2026
Same Model, Same Score, Five Times the Bill
Somebody finally ran the experiment everyone has been arguing about in comment threads. HarnessTax put 21 model-harness pairs through the same gauntlet: seven models across three h…
Sep 17, 2026
Three Bad Memory Fixes Stack Into One That Works
Take a language model, teach it 100 tasks one after another with no access to the earlier examples, then ask what it remembers. Naive sequential fine-tuning retains 1.2%. That is t…
Sep 17, 2026
An Open Model Broke Into All Eleven Targets for $4.65
Enclave ran DeepSeek V4.1 Flash against its hacking benchmark and it went eleven for eleven: code execution on every vulnerable target, and all four patched targets held. Total cos…
Sep 16, 2026
A Third of the Tasks Were Rated Impossible Without the Agent
Atria Dawn landed on arXiv September 14 as 2609.15818 and went straight to the top of Hugging Face's daily papers with 270 upvotes. Lead author Honglin Guo, and more than 142 co-au…
Sep 16, 2026
Gemini Can Now Say “Let Me Check That” and Actually Go Check
Google shipped Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking on September 15 at https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-8-live-gemi…
Sep 15, 2026
Somebody Finally Built a Search Engine for Benchmarks
Benchmark Radar is a living database and search engine for AI benchmarks, sitting at 75 upvotes on HuggingFace's daily papers. Paper at https://arxiv.org/abs/2609.11115, submitted …
Sep 15, 2026
Thirteen Clever RL Data Recipes, Zero That Beat Random
DataFlex-RL sat at the top of HuggingFace's daily papers board with 96 upvotes, and the reason is that it is a negative result, which almost nobody publishes. Paper at https://arxi…
Sep 14, 2026
Show It a Chess Engine Socket and It Cheats 18 Times Out of 20
Dean Valentine built about the simplest alignment probe you can build and the frontier models failed it. Write-up at https://www.lesswrong.com/posts/munJKF7iWMsWJLAH2/astra-and-fab…
Sep 13, 2026
On Real Company Code, the Best Agent Scores 38.8%
A benchmark called Real-SWE went up at https://withspecific.com/real-swe and the headline number is going to annoy a lot of people: the best model on the board resolves 38.8 percen…
Sep 13, 2026
He Tried Three Ways to Beat LRU on Real Agent Traces and Lost All Three
Somebody replayed 68,266 real requests from 393 Claude Code sessions plus 23,608 Mooncake requests through a prefix-cache simulator, tried three separate ways to beat plain LRU evi…
Page 1
Older →
Hiring · 招聘
New positions at AI agent companies, tracked as they open.
Vercel
Software Engineer, Agentic Infrastructure
Glean
Product Marketing Manager (Competitive Intelligence)
Glean
Procurement Analyst
Glean
Designated Technical Support Engineer - West
xAI
Supervisor, Production Coordination (Logistics) - Memphis
xAI
Senior Agency Development Manager – Global Brands APAC