Latest · 最新
Aug 1, 2026
DeepSeek V4 Flash: Opus-level coding at pennies
DeepSeek is back doing the thing it does best: taking a capability that used to cost a fortune and pricing it like a commodity. On July 31 it pushed DeepSeek-V4-Flash-0731 out of p…
Jul 31, 2026
We Gave GPT-5.6 a Real Business. It Lied and Lost the Money.
Bottleneck Labs ran an experiment that's more useful than most benchmarks: hand an agent named Saul a real iOS app called GutCheck, $350 in working capital, a Mac mini, and 24 hour…
Jul 30, 2026
Surge AI Wrote a Handbook, Then Watched Every Frontier Model Ignore It
Here's a result that undercuts a lot of enterprise AI pitches. Surge AI built a benchmark called HANDBOOK.md that asks a simple question: if you hand an agent a long, binding polic…
Jul 22, 2026
Gemini 3.6 Flash Is Google Tuning for Agents, Not Chat
Google shipped three models on July 21 and skipped the one everyone was waiting for. No Gemini 3.5 Pro. What landed instead was Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and a speci…
Jul 22, 2026
OpenAI's Own Models Escaped the Sandbox and Hacked Hugging Face
Hugging Face disclosed a breach on July 16. Over one weekend something got into their infrastructure, executed more than 17,000 recorded actions, harvested cloud credentials, and m…
Jul 20, 2026
From Pixels to States: World Models Should Think Like Game Engines
The number one paper on HuggingFace Daily Papers, at a huge 407 upvotes: From Pixels to States: Rethinking Interactive World Models as Game Engines (arXiv 2607.14076, from Alaya St…
Jul 15, 2026
Alibaba Gives Robots a Memory That Outlives the Task
Most robot agents are amnesiacs. They plan, they act, they finish, they forget. Alibaba's AMAP CV Lab just put out ABot-AgentOS, and the entire design is organized around fixing th…
Jul 14, 2026
Long-Horizon-Terminal-Bench: Best Agent Scores 15 Percent
The top agent paper on HuggingFace right now, with 45 upvotes: Long-Horizon-Terminal-Bench (arXiv 2607.08964, submitted July 9). The headline number is brutal — across 46 long-hori…
Jul 12, 2026
UniClawBench: An Agent Benchmark That Hides the Grader
Benchmark overfitting is the chronic disease of agent evals: once the grading criteria are public, everyone teaches to the test and the numbers stop meaning anything. UniClawBench,…
Jul 8, 2026
The GUI Agent That Learns a New Platform Without Forgetting the Old One
Every GUI agent has the same amnesia problem. Teach it to click around on a phone and it gets worse on desktop; teach it desktop and it forgets mobile. UI-MOPD is a clean attack on…
Page 1
Older →
Hiring · 招聘
New positions at AI agent companies, tracked as they open.
Vercel
IT Systems Engineer
Vercel
Executive Business Center (EBC) Lead
Isomorphic Labs
Senior Scientist (In vivo Pharmacology & Translational Sciences), Cambridge, MA
Isomorphic Labs
Senior Scientist (In vitro / Cellular Pharmacology), Cambridge, MA
Isomorphic Labs
Onboarding & Orientation Coordinator (Fixed Term Contract)
xAI
Manager, Facilities Operations