Staff Software Engineer, Agentic SDLC Foundations
- Location
- Onsite (San Jose, CA)
- Compensation
- $207k - $300k/yr
- Employment
- Full-time
- Level
- Senior Level
Posted 4 days ago
About the Role
Join Google's Core Engineering team to lead the technical strategy for building and scaling developer agent ecosystems. You will drive the architecture, evaluation, and quality engineering of agentic systems to transform how software is built and managed at scale.
Skills
Software Development
Software Architecture
Machine Learning
Model Deployment
Model Evaluation
Data Processing
Generative AI
LLM
Technical Leadership
Prompt Optimization
Fine-tuning
Agentic Systems
Quality Engineering
Benchmarking
Human-in-the-loop
System Design
Benefits
- Health Insurance
Perks
- Bonus
- Equity
Full job details
Minimum qualifications:
- Bachelor’s degree or equivalent practical experience.
- 8 years of experience in software development.
- 5 years of experience testing, and launching software products, and 3 years of experience with software design and architecture.
- 5 years of experience with ML design and ML infrastructure (e.g., model deployment, model evaluation, data processing, debugging, fine tuning).
- Experience integrating generative AI tools or LLM interfaces into workflows.
Preferred qualifications:
- Master’s degree or PhD in Engineering, Computer Science, or a related technical field.
- 3 years of experience in a technical leadership capacity (e.g., leading technical roadmaps, designing multi-system architectures, and guiding cross-functional engineering efforts).
- Experience designing and implementing AI/ML evaluations, LLM benchmarking frameworks, or agentic quality measurement pipelines.
- Experience in prompt optimization, LLM fine-tuning, automated dataset generation, and trajectory-based evaluation methodologies.
- Experience taking generative AI solutions or autonomous agent loops from prototype to high-reliability production systems.
About the job:
As part of the Agentic SDLC organization in Core, our mission is to transform how engineers build, deploy, and manage software at scale by creating an end-to-end agentic platform for the modern software development lifecycle.As a Staff Software Engineer on the Agentic SDLC Foundations team, you will serve as the technical lead driving the architecture, evaluation, and quality engineering of Google’s developer agent ecosystem. Bridging the gap between experimental generative AI prototypes and dependable, production-grade autonomous systems, you will lead the end-to-end technical strategy for building AI solutions with robust human-in-the-loop (HITL) guardrails. You will establish automated evaluation harnesses, benchmark suites, and methodologies that systematically elevate agent reliability, prevent regressions, and standardize key autonomy metrics across Google’s most critical engineering workflows.The Core team builds the technical foundation behind Google’s flagship products. We are owners and advocates for the underlying design elements, developer platforms, product components, and infrastructure at Google. These are the essential building blocks for excellent, safe, and coherent experiences for our users and drive the pace of innovation for every developer. We look across Google’s products to build central solutions, break down technical barriers and strengthen existing systems. As the Core team, we have a mandate and a unique opportunity to impact important technical decisions across the company.Individual pay is determined by factors including job-related skills, experience, and relevant education or training.
US: $207000 - $300000 (USD) + 20% bonus target + equity + benefits
Learn more about benefits at Google.
Responsibilities:
- Design and scale production developer agents, establishing patterns for trajectory planning, tool selection, and state recovery.
- Implement intelligent HITL approval gates, intent disambiguation, and graceful to ensure safe execution of actions.
- Architect evaluation pipelines that grade multi-step trajectories, tool-call fidelity, intermediate states, and reasoning paths beyond static input/output matching.
- Establish closed-loop pipelines converting execution failures and human corrections into golden datasets, prompt/tool-schema optimizations, and fine-tuning cycles.
- Curate end-to-end benchmarks replicating complex developer tasks to evaluate agent capabilities and establish standardized Key Performance Indicator (KPIs) across workflows.