Senior Software Engineer, Google Distributed Cloud AI
- Location
- Onsite (Sunnyvale, CA)
- Compensation
- $174k - $252k/yr
- Employment
- Full-time
- Level
- Senior Level
About the Role
Join Google's Distributed Cloud AI team to lead the technical design and development of software components for Large Language Model inference. You will drive integration across core services and optimize serving capabilities for hybrid cloud and edge environments.
Skills
Benefits
- Health Insurance
- Bonus
- Equity
Perks
- Bonus
- Equity
Full job details
Minimum qualifications:
- Bachelor’s degree or equivalent practical experience.
- 5 years of experience with software development in one or more programming languages.
- 3 years of experience testing, maintaining, or launching software products, and 1 year of experience with software design and architecture.
- 3 years of experience with developing large-scale infrastructure, distributed systems or networks, or experience with compute technologies, storage or hardware architecture.
- Experience programming in Go for software development, including AI/ML applications.
Preferred qualifications:
- Master's degree or PhD in Computer Science or related technical field.
- 5 years of experience with data structures and algorithms.
- 1 year of experience in a technical leadership role.
- Experience with container orchestration (e.g., Kubernetes) and cloud-based AI platforms.
About the job:
Google's software engineers develop the next-generation technologies that change how billions of users connect, explore, and interact with information and one another. Our products need to handle information at massive scale, and extend well beyond web search. We're looking for engineers who bring fresh ideas from all areas, including information retrieval, distributed computing, large-scale system design, networking and data storage, security, artificial intelligence, natural language processing, UI design and mobile; the list goes on and is growing every day. As a software engineer, you will work on a specific project critical to Google’s needs with opportunities to switch teams and projects as you and our fast-paced business grow and evolve. We need our engineers to be versatile, display leadership qualities and be enthusiastic to take on new problems across the full-stack as we continue to push technology forward.
If you are excited about shaping the future of AI on the edge and in hybrid clouds, and addressing unique challenges across various deployment models, we want to hear from you!
US: $174000 - $252000 (USD) + 15% bonus target + equity + benefits
Learn more about benefits at Google.
Responsibilities:
- Lead the technical design, development, and optimization of software components critical for Large Language Model (LLM) inference serving on GDC. This includes areas such as model life-cycle management, efficient data loading, dynamic request routing, and intelligent load balancing.
- Drive horizontal integration across core platform services, including billing, logging, observability, security, and quota management.
- Implement and enhance serving capabilities to support advanced LLM techniques like disaggregated serving, speculative decoding, quantization, and efficient model sharding across distributed hardware.
- Collaborate closely with internal teams developing core LLM frameworks, container orchestration (Kubernetes, Google Kubernetes Engine (GKE)), networking infrastructure, and hardware acceleration to build a cohesive and high-performance serving platform.