AIML - Software Engineer - AI, Evaluation
Apple
- Location
- Onsite (Cupertino, California)
- Employment
- Full-time
- Level
- Senior Level
Posted 2 days ago
About the Role
Apple's Evaluation organization is seeking an AI Software Engineer to design and build extensible frameworks and tools for assessing Apple AI features like Siri and Search. This role focuses on improving the quality and efficiency of model evaluations through principled assessments and cross-functional collaboration.
Skills
Python
Software Engineering
System Design
API Design
CI/CD
Testing Strategies
Debugging
LLM Applications
MLOps
Model Lifecycle Management
Data Leakage Analysis
Evaluation Metrics
AI Modeling
Framework Development
Pipeline Development
Full job details
Do you get excited by building software systems to enhance the automatic evaluation of various Apple AI products? Our Evaluation organization is responsible for providing principled assessments across a diverse range of Apple features, from Search, Siri to the latest Apple Intelligence capabilities.
Our team specializes in building LLM-as-judge and related tools to improve both the quality and efficiency of these evaluations. We are seeking a highly innovative and passionate AI software engineer to expand our tools and systems.
As an AI Software Engineer on the team, you will design and build tools and systems that sit at the intersection of AI modeling, software engineering, and product quality. You will design and develop extensible frameworks, pipelines, and tools that enable efficient development, deployment, and qualitative measurement of AI models. Due to the breadth of products supported, the role requires strong software design and engineering skills. Your work will directly influence product launch decisions and enable teams across Apple to iterate faster and with greater confidence.
BS/MS/PhD degree in Computer Science, Machine Learning, AI, or a related field. Exceptional Python skills. Solid software engineering fundamentals with production experience, including system design, API design, CI/CD, testing strategies, code maintainability, system monitoring, debugging complex systems and etc. Demonstrated expertise in using AI-assisted software development workflows to accelerate software development while maintaining code quality. Strong communication skills and proven ability to work collaboratively with cross-functional teams.
Experience with building LLM applications, frameworks, and offline evaluations. Familiar with MLOps principles for model lifecycle management. Experience in building scalable tools for product quality evaluation. Ability to understand and interpret evaluation reports, including metrics such as precision, recall, run-to-run consistency, and common pitfalls like data leakage. Product-minded, with a strong ability to translate ambiguous product requirements into solutions.
Description
As an AI Software Engineer on the team, you will design and build tools and systems that sit at the intersection of AI modeling, software engineering, and product quality. You will design and develop extensible frameworks, pipelines, and tools that enable efficient development, deployment, and qualitative measurement of AI models. Due to the breadth of products supported, the role requires strong software design and engineering skills. Your work will directly influence product launch decisions and enable teams across Apple to iterate faster and with greater confidence.
Minimum Qualifications
BS/MS/PhD degree in Computer Science, Machine Learning, AI, or a related field. Exceptional Python skills. Solid software engineering fundamentals with production experience, including system design, API design, CI/CD, testing strategies, code maintainability, system monitoring, debugging complex systems and etc. Demonstrated expertise in using AI-assisted software development workflows to accelerate software development while maintaining code quality. Strong communication skills and proven ability to work collaboratively with cross-functional teams.
Preferred Qualifications
Experience with building LLM applications, frameworks, and offline evaluations. Familiar with MLOps principles for model lifecycle management. Experience in building scalable tools for product quality evaluation. Ability to understand and interpret evaluation reports, including metrics such as precision, recall, run-to-run consistency, and common pitfalls like data leakage. Product-minded, with a strong ability to translate ambiguous product requirements into solutions.