Bio
Caspar is an Associate Professor at the University of Warwick. His recent work is primarily concerned with the measurement of AI welfare and sentience, as well as AI moral psychology. Previously he worked for several years on questions of statistical methodology for measuring human wellbeing, as well as more applied work on its determinants.
Projects
Research Direction: Inter-AI morality: How AI systems view and treat one another
AI systems increasingly cooperate, compete, and supervise one another. In ongoing work, we propose and begin the empirical study of “inter-AI morality”: how AIs view and treat other AIs, including whether they regard each other as deserving of moral consideration.
Building on this agenda, fellows will develop and run experiments with language models, analyse behavioural results, and, where appropriate, investigate and intervene on model internals. We are especially interested in projects combining white-box and black-box methods or going beyond dyadic interactions.
Possible directions include:
- Testing whether changing a model’s perception of another AI’s morally relevant properties changes its behaviour towards that system.
- Amplifying or suppressing inter-AI prosocial behaviour and measuring its wider effects on safety and alignment with human goals.
- Investigating how moral norms or obligations spread within groups of interacting AI systems.
Throughout we are especially interested in separating distinctly moral motivations towards other AIs from strategic considerations or obedience to hierarchy.
This work may matter for both human safety and for the potential welfare of future AI systems. Regarding the former, we really need to understand when inter-AI concern supports safe cooperation and when it undermines safeguards or human interests. Regarding the latter, if future AI systems can have welfare, how they treat one another will become the dominant determinant of welfare.
Fellows will have substantial scope to shape the specific research questions and methods within this broader agenda.
The aim for the end of the fellowship is to have a substantial set of empirical results and a paper draft suitable for development into a submission to a conference or general-interest journal.
What we’re looking for in a Mentee
We would work best with someone who enjoys turning broad conceptual questions into concrete experiments.
We also look for someone who is comfortable writing code and analysing data, including the responsible use of coding agents. Experience running experiments with language models would be valuable. Experience with model internals, activation steering, or multi-agent systems would be especially useful.
Applicants could come from machine learning, quantitative social science, psychology, philosophy, or another relevant background. Prior work on AI welfare is welcome but not necessary.
What we’re like as Mentors
- We see mentees as colleagues on a collaborative project and enjoy working together on both big ideas and the nitty-gritty of code, experimental design, and analysis.
- Meetings will primarily be online, with the option to meet in person at LISA in London or in Cambridge.
- Mentees will be integrated into the wider Cambridge Digital Minds community.
- The early part of the fellowship will focus on turning research questions from our forthcoming agenda paper into concrete empirical designs. Later stages will focus on running experiments, analysing results, and writing up the findings.


