Staff/Sr. Machine Learning Engineer, Foundation Models - AI, Search & Knowledge Platforms
Apple
- Location
- Onsite (Santa Clara, California ยท Seattle, Washington)
- Employment
- Full-time
- Level
- Senior Level
Posted 2 days ago
About the Role
Join Apple's Foundation Model Inference Team to build and optimize the inference stack powering Apple Intelligence and core services like Siri and Search. You will optimize large-scale language and vision models for low-latency, high-throughput performance serving billions of users globally.
Skills
Machine Learning
LLM Inference
CUDA
GPU Programming
Pytorch
Tensorflow
Kubernetes
Docker
Golang
Python
Transformers
Deep Learning
TensorRT-LLM
vLLM
DeepSpeed
Nvidia Triton Server
Full job details
We are Foundation Model Inference Team, within AI, Search & Knowledge Platform Technologies organization. Our team is responsible to build Inference stack to power Apple Intelligence. It builds frameworks, services and tools that power the largest Apple foundation models on servers. Our Infrastructure powers a wide gamut of services at Apple including Apple Search, Apple Music, AppleTV, AppStore, iMessages, Photos & Camera, Spotlight, Safari, Siri and upcoming ever exciting Apple products serving millions of queries every day with incredible low latencies, drawing every ounce of compute from our hardware. As part of this group, you will get a chance to bring Intelligence to billions of users across the world. You will have an opportunity to make difference in life of people by empowering them with AI. You will have a chance to work on optimizing billions of parameter language and vision and speech models using state of the art technologies and make it run at scale of Apple.
Work along side Foundation Model Research team to optimize inference for cutting edge model architectures. Work closely with product teams to build Production grade solutions to launch models serving millions of customers in real time. Build tools to understand bottlenecks in Inference for different hardwares and use cases. Mentor and guide engineers in the organization.
5+ years of experience leading and driving complex, ambiguous projects. Experience with LLM inference stack Familiarity with GPU programming concepts using CUDA. Familiarity with one of the popular ML Frameworks like Pytorch, Tensorflow. Have experience with high throughput services particularly at supercomputing scale. Proficient with running applications on Cloud (AWS / Azure or equivalent) using Kubernetes, Docker etc. Familiar with one of the popular ML Frameworks like Pytorch, Tensorflow. BS in Computer Science, Artificial Intelligence, Machine Learning, Information Retrieval, Data Science or related field
Proficient in building and maintaining systems written in modern languages (eg: Golang, Python) Familiar with fundamental Deep Learning architectures such as Transformers, Encoder/Decoder models. Familiarity with Nvidia TensorRT-LLM, vLLM, DeepSpeed, Nvidia Triton Server etc. Experience writing custom CUDA kernels using CUDA or OpenAI Triton. MS in Computer Science, Artificial Intelligence, Machine Learning, Information Retrieval, Data Science or related field.
Description
Work along side Foundation Model Research team to optimize inference for cutting edge model architectures. Work closely with product teams to build Production grade solutions to launch models serving millions of customers in real time. Build tools to understand bottlenecks in Inference for different hardwares and use cases. Mentor and guide engineers in the organization.
Minimum Qualifications
5+ years of experience leading and driving complex, ambiguous projects. Experience with LLM inference stack Familiarity with GPU programming concepts using CUDA. Familiarity with one of the popular ML Frameworks like Pytorch, Tensorflow. Have experience with high throughput services particularly at supercomputing scale. Proficient with running applications on Cloud (AWS / Azure or equivalent) using Kubernetes, Docker etc. Familiar with one of the popular ML Frameworks like Pytorch, Tensorflow. BS in Computer Science, Artificial Intelligence, Machine Learning, Information Retrieval, Data Science or related field
Preferred Qualifications
Proficient in building and maintaining systems written in modern languages (eg: Golang, Python) Familiar with fundamental Deep Learning architectures such as Transformers, Encoder/Decoder models. Familiarity with Nvidia TensorRT-LLM, vLLM, DeepSpeed, Nvidia Triton Server etc. Experience writing custom CUDA kernels using CUDA or OpenAI Triton. MS in Computer Science, Artificial Intelligence, Machine Learning, Information Retrieval, Data Science or related field.