
- Robustness & Security
- Personas & Model Organisms
- Evaluations
Sampura Research Stream
Bio
Josh is the co-founder and CTO of Sampura Research. Previously, he co-led the Human Data Engineering organization at Google DeepMind, helping curate high quality human and AI data for use in model training and evaluation across a wide variety of capability and safety research.
Projects
Research Direction: Evaluating the robustness of frontier judges
At Sampura Research, we are building better human-AI judges to reduce reward hacking and other harmful behaviors which emerge during model training. We are interested in mentoring two types of projects to further research around measuring and improving these judges:
- Dynamic adversarial datasets: research judge performance under optimization pressure
a. Our current leaderboard evaluates a judge's static accuracy on a task.
b. At training time, judges will be subject to optimization pressure as the reward policy updates to satisfy a given judge.
c. We would like to explore how our current and future judge methods will perform in a training time scenario, including:
i. Building infrastructure to run RL with different judge protocols
ii. Creating/curating datasets which would be most useful to evaluate dynamically
iii. Define and evaluate good metrics for determining judge “performance” during RL
iv. Evaluate the behavior of “co-training” parts of the judge (e.g. judge policy gets updated iteratively along with the policy under training based on outcomes).
- Create SOTA datasets which fill in “holes” in our current static evalset for judges
a. Axis:
i. Domain: Complex safety related domains such as sociotechnical harms, as well as alignment tasks like monitoring / scheming.
ii. Interaction Type: e.g. Agentic, multi-turn, long-context
b. This will likely involve working directly with synthetic data pipelines, human data vendors and/or subject matter experts to procure both high quality tasks and high quality ground truth data.
For both areas, we have specific projects in mind, but are also open to discussing other directions/projects which could help accomplish these top-level goals. A strong result for either of these projects would be a meaningful blog post or leaderboard release, providing insight on how our judges perform under a diverse set of tasks or direct optimization pressure. This work directly amplifies our research agenda: https://sampura.org/news/announcing-sampura-research/
What we’re looking for in a Mentee
We work best with fellows who are largely self motivated, take the time to understand the broader research goals, and are willing to get their hands dirty in a wide variety of areas (e.g. pure research, small infra tweaks) in order to maximize impact. We are open to fellows of any career stage or background, but generally value a demonstrated history of strong research output and engineering rigor.
What we’re like as Mentors
As a mentor, I really value in-person and ad-hoc communication. We'll be based in London, and will add mentees to our team Slack, making it very easy to have individual and group ad-hoc discussions. For recurring meetings, I tend to prefer short and focused time to discuss any blockers or ambiguities, rather than focus on status updates which can be provided async. I expect to be more hands on during the start of the program as we decide on the final project goals, and more hands off once fellows get up to speed.