Projects

Explore the projects our fellows work on with their mentors, cohort by cohort.

Q3 2026

AI Safety Fellowship

Explore the projects, fellows, and mentors of Pivotal's 2026 Q3 AI Safety Research Fellowship.

  • Ongoing
  • AI Governance

A benchmark for AI agent compliance with US sanctions law

A benchmark evaluating whether AI agents comply with US sanctions law.

Fellow:
Arjun Sharma
Mentor:
Kevin Wei
  • Ongoing
  • Technical AI Safety

A benchmark for “deal-making” with potentially scheming models

Building a benchmark that probes whether and how models engage in "deal-making" — offering a potentially-scheming model something in exchange for revealing misalignment — with a focus on qualitatively analysing the model's reasoning about credibility, cheating and being evaluated.

Fellow:
Mark Keavney
  • Ongoing
  • AI Governance

A disclosure framework for AI-lab insiders

Developing a framework to help AI-lab insiders judge whether and when to disclose high-impact information.

  • Ongoing
  • Technical AI Safety

A simpler “cross-domain solution” for protecting model weights

Designing and prototyping a simpler, smaller-attack-surface "cross-domain solution" for securely moving data — the kind of security-critical infrastructure needed to protect model weights against sophisticated state-level attackers under the SL5 standard.

Mentor:
Guy
  • Ongoing
  • Technical AI Safety

A taxonomy for compositional interpretability and hierarchical concepts

Working toward a taxonomy for compositional interpretability and hierarchical concepts — top-down interpretability techniques aimed at pragmatic alignment goals.

Fellow:
Ethan Nguyen
  • Ongoing
  • AI Governance

A threat model for Autonomous Replication and Adaptation (ARA)

Producing a rigorous threat model and capability-threshold analysis for Autonomous Replication and Adaptation (ARA) — working out where the weakest links are in the causal chain, and whether current lab safeguards adequately cover the threat.

Mentor:
Jide Alaga
  • Ongoing
  • Biodefense

An agent that refreshes biology-experiment evals from recent publications

Recent evals have begun to go beyond human-expert capability by asking models to predict real-world biology experiment outcomes — but that needs a constantly refreshed pool of experiments outside the training corpus to avoid measuring memorisation. Kamal is building an agent that automatically refreshes that pool by finding and formatting experiments from recent publications and preprints as eval tasks.

Fellow:
Kamal Maher
  • Ongoing
  • Biodefense

Applying Item Response Theory to item-wise eval data

Applying Item Response Theory — the method behind the Epoch Capabilities Index — to item-wise eval data at SecureBio, to find the most information-rich items.

Fellow:
Ben Pomeranz
  • Ongoing
  • Technical AI Safety

Auditing Bloom as a measurement instrument for evaluating language models

Auditing the effectiveness of Bloom as a measurement instrument for evaluating language models.

Fellow:
Ben Slater
Mentor:
Kevin Wei
  • Ongoing
  • Technical AI Safety

Can LLM agents model themselves and other agents?

Testing whether LLM agents can model themselves and other agents in multi-agent settings — and whether, when a model fails to predict another agent, its guess gets pulled toward what it itself would have done. Relevant to any setting where one model monitors another.

Fellow:
Luis Montoya
  • Ongoing
  • Technical AI Safety

Challenging LLM self-introspection techniques from an adversarial standpoint

Improving and challenging current LLM self-introspection techniques from an adversarial standpoint.

Fellow:
Guy Galun
Mentor:
Peter Hase
  • Ongoing
  • Technical AI Safety

Character traits, reward hacking, and chain-of-thought monitorability

Studying how character traits shift reward-hacking rates and chain-of-thought monitorability, and whether a teacher model's "corrupted" self-account transmits to a student distilled on its reasoning.

Fellow:
Ionuț Stan
  • Ongoing
  • Technical AI Safety

Deception detectors as reward signals for honesty training

Using deception detectors as reward signals for honesty training, developing methods to reduce obfuscation.

Fellow:
Lily Shi
Mentor:
Peter Hase
  • Ongoing
  • Technical AI Safety

Extending AI control protocols with interrogation

Building a control protocol that uses interrogation: a trusted model asks an untrusted model pointed questions about each proposed action, and a trusted judge uses the answers to catch dangerous actions while keeping useful ones flowing. The idea is that convincing justifications are easier to give when there’s a genuine reason, which should make the monitor more effective.

Mentor:
Adam Kaufman
  • Ongoing
  • Technical AI Safety

How agents influence one another over long-horizon interactions

Investigating how agents influence one another over long-horizon, multi-turn interactions and tasks.

Fellow:
Ben Maltbie
  • Ongoing
  • Technical AI Safety

How framing persona evidence as memory shapes model behaviour

Testing whether the same persona evidence shapes a model more strongly when framed as its own prior behaviour or memory than when framed as another assistant's output or an external document — with implications for how memory, summaries and handoffs should be designed.

  • Ongoing
  • Technical AI Safety

Interpretability for monitoring and steering LLMs

Working in Peter Hase's stream on interpretability for monitoring and steering LLMs — scoping a concrete project that turns interpretability into something practical for oversight.

Fellow:
Thea Xu
Mentor:
Peter Hase
  • Ongoing
  • AI Governance

Mapping extreme power concentration across frontier labs and governments

Mapping extreme power concentration across frontier labs and governments: where the power actually lies, what kind of power it is, and how reversible it is.

Fellow:
Haimi Tefera
  • Ongoing
  • AI Governance

Mapping the U.S. response function for a middle-power AI coalition

Mapping the full U.S. response function — from active partnership and export exemptions through to restrictive instruments such as export controls, sanctions and supply-chain pressure — to identify the design choices that most affect the viability of a middle-power frontier-AI coalition, drawing on historical examples.

  • Ongoing
  • Technical AI Safety

Mapping the “attractor” states LLMs drift into over long interactions

Investigating the "attractor" states LLMs drift into over long interactions (Claude's spiritual-bliss basin, Gemma's frustration spiral). He'll run many fast experiments to find basins systematically and see how character training, memory and compaction reshape them.

Fellow:
Tim Farrelly
  • Ongoing
  • AI Governance

Middle-power interests and prospective AI coalitions

Mapping the interests of middle-power nations and the kinds of coalitions they'd be interested in forming.

  • AI Governance

Physical side-channels as verification layers for international AI agreements

Developing verification layers for international AI agreements: researching physical side-channels (electromagnetic and power emissions) as fingerprints for LLM workload verification, letting a verifier classify what a cluster is actually running without trusting the operator.

Mentor:
Gabriel Kulp
  • Ongoing
  • Technical AI Safety

Scientific reliability when alignment research is automated

In Kozzy Voudouris's stream, modelling what happens to scientific reliability when alignment research itself is automated.

  • Ongoing
  • AI Governance

Should more people be working on AI-enabled power concentration?

Should more people be working on AI-enabled power concentration?

Fellow:
Hugo Bos
  • Ongoing
  • Technical AI Safety

Tensor-network mechanistic interpretability for biological foundation models

Tensor-network mechanistic interpretability for biological foundation models.

  • Ongoing
  • AI Governance

The scheming threat landscape in national-security AI deployment

Mapping the scheming threat landscape within national-security and military AI deployment — which pathways are most plausible, most dangerous and least addressed, and whether scheming there is more likely to come from external interference or internal misalignment.

Fellow:
Liha Leung
Mentor:
Jide Alaga
  • Ongoing
  • Technical AI Safety

What causes activation plateaus in models?

Investigating what causes “activation plateaus” in models — testing whether they arise because activations live on a non-linear manifold, and whether perturbing along the manifold versus across it predicts where the plateaus and their sensitive directions appear. This could give a data-independent way to discover a model’s features and steer it.

  • Ongoing
  • Technical AI Safety

Why models become evaluation-aware during post-training

Understanding why models become evaluation-aware during post-training, and whether there are robust interventions for mitigating the rise in evaluation awareness and metagaming. Ryan is also collaborating with Ben Slater on the Bloom auditing project.

Mentor:
Kevin Wei

Get involved

Join our team

Work with us directly to advance AI safety research. No roles are open at the moment. If you think you may be a good fit for the team, introduce yourself and we will reach out when something opens up.

Mentor top emerging researchers

Guide a fellow on a research project, share your expertise, and help them grow as a researcher. We handle recruitment, logistics, and day-to-day research support.

Participate in a fellowship

Join our in-person research fellowship in London and work with experienced mentors on AI safety, AI governance, AIxBio, or biodefense.