Latest · 最新
Sep 18, 2026
ScienceIDE Turns Scientific Repos Into Places an Agent Can Actually Learn
A paper from PhAI Labs, submitted September 16 and sitting near the top of Hugging Face daily papers, names a problem the agent field has been dancing around: the scientific experi…
Sep 18, 2026
Xiaomi Is Streaming Its RL Run Live and Nobody Else Would Dare
Luo Fuli, who runs Xiaomi's MiMo team and came out of DeepSeek, broke roughly six months of silence on September 16 by doing the single most un-lab-like thing available: turning th…
Sep 17, 2026
A 4B Model Beat the Postgres Query Planner for $1,200
Postgres has had thirty years of very smart people tuning its query planner. Rohan Bansal trained a 4B model to write hints that beat it by 1.81x geometric mean on the Join Order B…
Sep 17, 2026
Dream-RSI Improves Itself by Replaying Its Own Failures
The expensive part of a self-improving agent is not the improving, it is the trying. Every policy change has to be tested against the real world, and in algorithm engineering or GP…
Sep 15, 2026
Thirteen Clever RL Data Recipes, Zero That Beat Random
DataFlex-RL sat at the top of HuggingFace's daily papers board with 96 upvotes, and the reason is that it is a negative result, which almost nobody publishes. Paper at https://arxi…
Sep 10, 2026
Miles v0.1 Is the Post-Training Stack Labs Usually Don't Open Source
RadixArk dropped Miles v0.1 and called it production-level post-training, which is a boring name for the least boring thing on Hugging Face's board today. This is the plumbing unde…
Sep 10, 2026
NeoHorse-1 Ships the Loop Everyone Says They're Afraid Of
While a researcher was quitting Anthropic over self-improving AI, a team called TokenRhythm put a working prototype of it on Hugging Face under Apache 2.0. NeoHorse-1 is two checkp…
Sep 6, 2026
Two Papers, Same Day, Same Message: Agent RL Is Out of Environments
Two papers landed on arXiv on September 3 saying the same thing from opposite directions: the bottleneck in agent RL is no longer models or compute, it is environments to train the…
Sep 1, 2026
J-Zero Grows Challenger, Solver and Judge From Zero Data
Self-improvement loops have a standard cause of death: the judge. Fix a reward model or an LLM judge in place, let the policy improve against it, and within a couple of rounds the …
Aug 29, 2026
Agents Build Games So World Models Have Something to Learn From
The top paper on Hugging Face's daily board (118 upvotes) inverts the usual relationship between agents and environments. Instead of training agents inside games, "Agentic Game Dev…
Page 1
Older →
Hiring · 招聘
New positions at AI agent companies, tracked as they open.
Vercel
Software Engineer, Agentic Infrastructure
Glean
Product Marketing Manager (Competitive Intelligence)
Glean
Procurement Analyst
Glean
Designated Technical Support Engineer - West
xAI
Supervisor, Production Coordination (Logistics) - Memphis
xAI
Senior Agency Development Manager – Global Brands APAC