Senior Machine Learning Engineer, Apple Cloud AI
Apple
- Location
- Onsite (Seattle, Washington)
- Employment
- Full-time
- Level
- Mid Level
Posted 6 days ago
About the Role
Apple's Apple Services Engineering organization is seeking a Senior Machine Learning Engineer to build and optimize managed platform services for frontier AI. The role focuses on the full AI lifecycle, from data engineering to model deployment, making AI models faster, cheaper, and more efficient for internal teams.
Skills
Python
Rust
Java
Machine Learning
Distributed Systems
Data Processing
Model Serving
Inference Optimization
LLM
Feature Engineering
Ray
Kubernetes
Cloud Infrastructure
API Development
ML Pipelines
Governance
Full job details
Apple is a place where extraordinary people gather to do their best work. Together we build products and experiences people love. The Apple Services Engineering (ASE) organization builds and operates the systems and infrastructure that power Apple's services at scale.
The Apple AI platform within ASE enables teams across Apple to build, train, optimize, and deploy AI systems at scale. Our team builds the optimization and intelligence layer for frontier AI, making frontier class of models work better, cheaper, and faster through managed, serverless capabilities that span the full AI lifecycle: data and feature engineering, embeddings and retrieval, model training and fine-tuning, inference optimization and routing, prompt optimization, evaluation, and governance.
We are looking for an ML engineer who is excited about building managed platform services at the intersection of ML, distributed systems, and production engineering.
3+ years of experience building production ML systems or ML infrastructure Strong programming skills in Python and/or Rust/Java Understanding of end-to-end machine learning workflows - from data preparation through training, evaluation, and deployment Experience with distributed systems and large-scale data processing Experience with model serving, inference optimization, or ML pipeline engineering Experience building APIs and services that other engineers consume Strong collaboration and communication skills Comfortable navigating ambiguity in fast-moving areas BS, MS, or PhD in Computer Science or equivalent practical experience
Experience with LLM inference optimization (batching, quantization, KV caching, tensor parallelism) Experience with model serving frameworks (vLLM, TensorRT, Ray Serve, or similar) Experience with embedding models and retrieval systems - fine-tuning encoders on graded or contrastive objectives, pooling strategies, dimensionality reduction for serving cost, vector databases, and retrieval evaluation (NDCG, recall, graded relevance) Experience with fine-tuning and alignment workflows (SFT, DPO, LoRA, RLHF, RLVR, GRPO, reward modeling) Experience with feature engineering and feature serving platforms (e.g. Feast, Tecton, Hopsworks), distributed data processing frameworks (e.g. Spark, Flink, Ray), offline stores (e.g. Iceberg, Delta, Lance), and online stores (e.g. Redis, Cassandra, DynamoDB) Experience with Ray, Kubernetes, and cloud GPU infrastructure (AWS, GCP) Experience with ML governance, lineage, or compliance systems
Description
We are looking for an ML engineer who is excited about building managed platform services at the intersection of ML, distributed systems, and production engineering.
Minimum Qualifications
3+ years of experience building production ML systems or ML infrastructure Strong programming skills in Python and/or Rust/Java Understanding of end-to-end machine learning workflows - from data preparation through training, evaluation, and deployment Experience with distributed systems and large-scale data processing Experience with model serving, inference optimization, or ML pipeline engineering Experience building APIs and services that other engineers consume Strong collaboration and communication skills Comfortable navigating ambiguity in fast-moving areas BS, MS, or PhD in Computer Science or equivalent practical experience
Preferred Qualifications
Experience with LLM inference optimization (batching, quantization, KV caching, tensor parallelism) Experience with model serving frameworks (vLLM, TensorRT, Ray Serve, or similar) Experience with embedding models and retrieval systems - fine-tuning encoders on graded or contrastive objectives, pooling strategies, dimensionality reduction for serving cost, vector databases, and retrieval evaluation (NDCG, recall, graded relevance) Experience with fine-tuning and alignment workflows (SFT, DPO, LoRA, RLHF, RLVR, GRPO, reward modeling) Experience with feature engineering and feature serving platforms (e.g. Feast, Tecton, Hopsworks), distributed data processing frameworks (e.g. Spark, Flink, Ray), offline stores (e.g. Iceberg, Delta, Lance), and online stores (e.g. Redis, Cassandra, DynamoDB) Experience with Ray, Kubernetes, and cloud GPU infrastructure (AWS, GCP) Experience with ML governance, lineage, or compliance systems