Expert Research Expert-Bench
DeepMind
Objective:
Develop a new version of the RE-Bench evaluation benchmark (used
for hillclimbing by Code Strike and others) by collecting high-quality, complex,
"hard" tasks in a red-teaming setup.
Working Model:
Tight feedback loop with researchers, including daily syncs.
Curators will have access to internal tools (Jetski, google3 codebase) and
unreleased model checkpoints.
Summary:
The Research Engineer (RE) Curator will design, implement, and review complex, multistep ("mid-horizon") agentic tasks that simulate real-world challenges faced by
Research Engineers. Each task will require 1-2 days of continuous effort to complete
and will span multiple technical skills.
Typical Work Activities (Examples)
Asked to create tasks such as:
Model Training Experimentation: Given a vague research idea (e.g., "modify
how RL reward is computed"), implement the change, run the training
experiment and analyze the results to determine success.
Methodology Comparison: Compare two different anomaly detection
algorithms on a dataset, calculate correlation, perform manual spot checks, and
summarize the findings in a Colab notebook for researcher decision-making
· Computer Science & Software Engineering (Core) Public
Python Scripting: Proficiency in writing and debugging Python code.
Development Infrastructure: Familiarity with version control systems (Git),
Integrated Development Environments (IDEs), and basic software development
workflows.
· Agentic Coding: Conceptual understanding or experience using AI coding assistants
(Gemini, Jetski) and understanding prompt engineering or agent workflows. Able to clean code practices, readability, and basic debugging skills.
· Machine Learning & Artificial Intelligence (Core)
ML Theory & Practice: Foundational understanding of machine learning concepts,
model training, and evaluation.
Large Language Models (LLMs): Familiarity with LLM capabilities, limitations, and
evaluation techniques.
Reinforcement Learning (RL): Basic understanding (helpful for task design involving
reward functions).
ML Experimentation: Experience setting up, running, and analyzing ML experiments.
· Data Science & Quantitative Analysis (Core)
Data Analysis & Analytics: Heavy data analysis skills, including statistical
correlation, data cleaning, and interpretation.
Notebook Environments: Proficiency in using Jupyter Notebooks or Google Colab
for analysis and reporting.
· STEM Research & Experimental Methodology (Core)
Strong background in experimental design, hypothesis testing,
and rigorous evaluation.
Computational Fields: Experience in STEM fields or Computational
Humanities/Social Sciences requiring significant computational work.
· Quality Assurance & Testing (Preferred / Plus)
Test Engineering experience designing test cases, quality review processes, and
debugging complex systems.
· AI Safety & Security (Preferred / Plus)
LLM Red Teaming: Experience in identifying vulnerabilities, edge cases, or failure
modes in LLMs.
Must-Haves:
research-heavy domain requiring data analysis and coding (e.g., computational
sociology/humanities).
ability to work independently and handle ambiguity.
Nice-to-Haves:
Pay: $150.00 - $160.00 per hour
Application Question(s):
Education:
Experience:
Language:
Work Location: Remote
Sign in to browse authentic reviews, anonymous ratings and salary data before you apply.