Latest · 最新
Jul 31, 2026
SkillRise: Agents That Grow Skills Instead of Relearning
A new paper called SkillRise (arXiv 2607.26784) goes after one of the more honest weaknesses of today's agents: they relearn everything from scratch. Solve a task, throw away what …
Jul 28, 2026
NVIDIA Molt: an RL framework small enough for the agent to read itself
NVIDIA's NeMo team dropped Molt today and it shot to the top of Hugging Face papers with 647 upvotes. It's a PyTorch-native training framework for agentic reinforcement learning, a…
Jul 22, 2026
DeepSearch-World: A 9B Search Agent That Taught Itself, No Teacher
The default recipe for a good search agent right now is: take a small model, distill trajectories out of a frontier model, ship it. HKUST's DeepSearch-World, arXiv 2607.07820, does…
Jul 19, 2026
LongStraw: everyone serves 1M context, almost nobody can train on it
Here's an embarrassing asymmetry in the agent stack: inference systems happily serve million-token contexts, but RL post-training — the thing that actually makes agents better at l…
Jul 18, 2026
SEED teaches an agent by making it its own tutor
Long-horizon agents trained with RL have a dumb problem: you only find out at the very end whether the whole episode worked, so the model gets almost no signal about which of its h…
Jul 11, 2026
Prime Intellect Raises $130M Series A So Enterprises Can Train Their Own Agents
Prime Intellect announced a 130 million dollar Series A at a 1 billion dollar valuation on July 8, led by Radical Ventures with Nvidia Ventures, Intel Capital, Dell Technologies Ca…
Jul 10, 2026
SAO: the RL trick that keeps long-horizon agents from blowing up
Everyone wants agents that run for hundreds of steps. Almost nobody can train them without the reinforcement learning going unstable. A new paper from the Zhipu/GLM team — Single-R…
Jul 9, 2026
Cognition's SWE-1.7 bets the ceiling is fake
Cognition dropped SWE-1.7 today, the coding model that runs inside Devin. The headline claim is not that it beats everyone, it doesn't. On FrontierCode 1.1 it hits 42.3% against GP…
Jul 7, 2026
This paper says everyone's been optimizing the wrong policy in LLM RL
A team out of Alibaba's Taobao group put out a paper with a blunt title: The Mirage of Optimizing Training Policies. The claim is that in RL for LLMs, we've been optimizing the wro…
Jul 6, 2026
EvoPolicyGym asks the real question: not can the agent solve it, but can it get better
Most benchmarks score the final answer. EvoPolicyGym scores the climb. It hands an agent a chunk of Python policy code, 16 environments, and a fixed 128-episode interaction budget,…
Page 1
Older →
Hiring · 招聘
New positions at AI agent companies, tracked as they open.
Vercel
IT Systems Engineer
Vercel
Executive Business Center (EBC) Lead
Isomorphic Labs
Senior Scientist (In vivo Pharmacology & Translational Sciences), Cambridge, MA
Isomorphic Labs
Senior Scientist (In vitro / Cellular Pharmacology), Cambridge, MA
Isomorphic Labs
Onboarding & Orientation Coordinator (Fixed Term Contract)
xAI
Manager, Facilities Operations