Senior Research Engineer, Agentic Data and Tooling, DeepMind
- Location
- Onsite (New York, NY)
- Compensation
- $174k - $252k/yr
- Employment
- Full-time
- Level
- Senior Level
About the Role
Join Google DeepMind's Agent Data and Tooling team to build high-velocity agentic data infrastructure and interactive environments that power Gemini's frontier capabilities. You will bridge research and engineering to drive model performance through curated data integration and gold-standard benchmark evaluation.
Skills
Benefits
- Health Insurance
Perks
- Bonus
- Equity
Full job details
Minimum qualifications:
- Bachelor's degree in Computer Science, Information Technology, a related technical field, or equivalent practical experience.
- 5 years of experience working with large language models (LLMs).
- 2 years of experience developing and training machine learning models.
- Experience with Agentic integrations and Model Context Protocol.
Preferred qualifications:
- Master's degree or PhD in Electrical Engineering, Computer Science, or equivalent practical experience.
- 2 years of experience with full-stack development.
- Excellent analytical, problem-solving and communication skills with demonstrated attention to detail.
- A deep passion for AI technology and all of its possibilities .
About the job:
At Google, research-focused Software Engineers are embedded throughout the company, allowing them to setup large-scale tests and deploy promising ideas quickly and broadly. Ideas may come from internal projects as well as from collaborations with research programs at partner universities and technical institutes all over the world.
From creating experiments and prototyping implementations to designing new architectures, engineers work on real-world problems including artificial intelligence, data mining, natural language processing, hardware and software performance analysis, improving compilers for mobile platforms, as well as core search and much more. But you stay connected to your research roots as an active contributor to the wider research community by partnering with universities and publishing papers.
The Google Deepmind (GDM) Agent Data and Tooling team within the Human Data Platform organization builds the environments, pipelines, and tooling that power Gemini's frontier agentic capabilities. We operate under an active data ownership model — moving beyond commodity data collection to build high-fidelity interactive worlds, capture complex multi-turn trajectories, and land data into model training and evaluation pipelines to drive measurable hillclimbing on coding, computer control, tool-use benchmarks, and more.
Build the critical infrastructure and interactive environments that directly drive Gemini's agentic and reasoning capabilities.
In this role, you will sit at the intersection of software engineering and model training: writing high-velocity production code to create rich interactive worlds.
We are looking for an engineer who loves to deliver code and build 0-1 systems at lightning pace and under high pressure, making heavy use of AI tools to boost velocity/output.
Artificial intelligence will be one of humanity’s most transformative inventions. At Google DeepMind, we are a pioneering AI lab with exceptional interdisciplinary teams focused on advancing AI development to solve complex global challenges and accelerate high-quality product innovation for billions of users. We use our technologies for widespread public benefit and scientific discovery, ensuring safety and ethics are always our highest priority.
US: $174000 - $252000 (USD) + 15% bonus target + equity + benefits
Learn more about benefits at Google.
Responsibilities:
- Build agentic data infrastructure at high velocity: Own and deliver key components across the agentic data stack. Rapidly prototype, iterate, and ship robust production code to generate, capture, and curate complex multi-turn agent trajectories at scale.
- Bridge research and engineering to drive model hillclimbing: Collaborate closely with Gemini research teams to close the loop between data creation and model quality. Understand how multi-turn trajectory design, environment complexity, and reward signals impact SFT, RL training, and capability hillclimbing. Directly integrate curated data into training pipelines and evaluate downstream model performance on frontier benchmarks.
- Build quality tooling and support gold-standard benchmarks: Create human-in-the-loop annotation tooling and interactive trajectory review surfaces, working in tandem with automated validation checkers leveraging adversarial LLM judges and programmatic verifiers. Support the creation and curation of gold-standard evaluation sets for flagship benchmarks.