Frontier AI agent research
Agents that reason
over long horizons.
xMAD.ai is a research company advancing large language model technology. We study how models plan, use tools, and stay reliable across thousands of steps — and we publish what we learn.
Research collaborations & compute partners
- Meridian Labs
- Northgate Institute
- Helio Compute
- Cambridge ALIGN
- Vector Foundry
- Ostro University
Research agenda
Four problems we work on.
Our agenda is narrow on purpose. Every project either extends how far an agent can act without supervision, or makes its behaviour easier to measure and trust.
-
01
Long-horizon reasoning
Keeping an agent coherent across tens of thousands of tokens of state: hierarchical planning, memory that survives context rollover, and recovery from its own mistakes.
-
02
Tool use & embodied control
Grounding a model in real interfaces — browsers, terminals, APIs, robots — and training it to know when a tool is the wrong one.
-
03
Evaluation & interpretability
Benchmarks that resist saturation, sparse autoencoders over agent trajectories, and methods for auditing a decision after the fact.
-
04
Safe deployment
Capability thresholds, sandboxed execution, and monitoring that catches a drifting agent before it causes harm rather than after.
Publications
Recent work.
-
Sierra: an agent that recovers from its own failures
We train a recovery policy on twenty million annotated failure trajectories. On a suite of unattended data-engineering tasks, Sierra completes 71.4% end-to-end, against 43.2% for a comparable agent without recovery training.
-
Tessera: measuring what an evaluation set no longer tests
Contamination and saturation are usually reported after the fact. We give a cheap online estimator for the marginal information a benchmark still carries about a model family, and show it flags seven widely used suites as exhausted.
-
Harbour: sparse features of agent trajectories
Applying dictionary learning to full agent trajectories rather than single activations yields features that transfer between tasks. We use them to predict, with 0.81 AUC, whether a rollout will end in a specification violation.
Models
Weights we release.
Four models are available under a permissive licence. Frontier checkpoints are shared with research partners under a capability agreement.
| Model | Parameters | Context | AgentBench-2 | Licence |
|---|---|---|---|---|
| xMAD-2 Mini | 3B | 128k | 48.1 | Apache 2.0 |
| xMAD-2 Base | 9B | 256k | 61.7 | Apache 2.0 |
| xMAD-2 Pro | 52B | 512k | 72.9 | xMAD Research |
| xMAD-2 Pro Long | 52B | 2M | 74.3 | xMAD Research |
AgentBench-2 scores are pass@1 on the held-out split, averaged over three seeds. Full methodology in the model card.
Safety
Written down, not implied.
Every model we train passes a documented review before release. The review covers capability thresholds, evaluation on autonomy-relevant tasks, and a red-team pass with external reviewers. Findings are summarised in the model card, including the ones that went badly.
- Capability thresholds published before training begins
- Independent red-teaming for every frontier checkpoint
- Incident reports published within 30 days
- Sandboxed execution by default in all hosted agents
“The interesting question is no longer what a model knows. It is what a model does on the two-hundredth step, when nobody is watching.”
Careers
We are hiring across four teams.
Research notes, once a month.
New papers, evaluation results, and occasional negative results. No product announcements.
Thanks — please confirm via the email we just sent.