xMAD.ai

Company

A research company, not a product company.

xMAD.ai was founded in 2023 to work on one question: what happens when language models stop answering and start acting. Everything else follows from that.

About

We are 140 researchers, engineers and operators across three offices. About two thirds of the company works on research; the rest builds the infrastructure that makes the research reproducible — evaluation harnesses, training pipelines, and the sandboxes our agents run in.

The company is funded by research partnerships rather than by a consumer product. That choice is deliberate: it keeps our incentives aligned with publishing results instead of shipping features on a quarterly cadence.

How we work

  • Research results are published, including negative results.
  • Evaluation harnesses ship with the paper, not after it.
  • No capability claim without a released evaluation protocol.
  • Every frontier training run has a named safety reviewer outside the team.

Safety policy

Our policy has three parts: thresholds defined before training begins, an independent review before any release, and public incident reports within thirty days of a confirmed issue.

Thresholds cover autonomy-relevant capabilities — long-horizon self-direction, irreversible actions, and deception under evaluation. Where a threshold is approached but not crossed, we say so in the model card and publish the measurements.

Questions about the policy are welcome from researchers and regulators alike.

Team

The leadership team, with research staff listed in full on request.

Dr. Imogen Hartley

Chief Scientist

Previously led evaluation research at a national AI institute. Works on capability thresholds and measurement.

Rafael Okafor

Head of Research

Long-horizon planning and recovery. Author of the Sierra line of work.

Dr. Lena Tanaka

Head of Interpretability

Sparse methods for sequential state. Previously at a university ML group in Zürich.

Priya Bhatt

Head of Engineering

Training and inference infrastructure. Built the sandbox our hosted agents run in.

Dr. Tomas Volkov

Safety Lead

Independent reviewer for frontier releases. Background in formal methods.

Marguerite Ferreira

Head of Evaluation

Benchmark design and decay estimation. Maintains AgentBench-2.

Dr. Yusuf Adeyemi

Research Scientist

Trajectory interpretability and violation prediction.

Sofia Lindqvist

Chief Operating Officer

Research partnerships, compute procurement and publishing.

Press kit

Logos, the wordmark, team photography and our boilerplate are available for editorial use. For interviews, contact the communications desk and we will respond within two business days.

Contact communications