Skip to content
AI Engineer Jobs
Onsite (San Francisco, California) $200k - $320k/yr Full-time Senior Level 3 benefits + 6 perks
Posted 2 days ago

About the role

Build AI is scaling physical labor datasets by co-designing hardware, infrastructure, and research. This role focuses on optimizing inference performance to reduce compute costs and enable scalable model deployment.

Skills

Machine learning Inference optimization Python C++ Rust CUDA PyTorch JAX Profiling Quantization Compilers TensorRT MLIR TVM Systems engineering
Remote (Seattle, Washington) $140k - $225k/yr Full-time Senior Level 2 benefits + 3 perks
Posted 2 days ago

About the role

AZX accelerates positive impact in critical industries like energy and real estate through physics-informed ML and enterprise AI solutions. This role owns the AI inference and runtime platform, managing Kubernetes infrastructure, model serving, and isolation boundaries to enable scalable, secure model deployment.

Skills

Rust Kubernetes Python FastAPI LLM serving vLLM SGLang Infrastructure as code Terraform OpenTofu Distributed systems GPU scheduling Isolation technology Observability OpenTelemetry System architecture
Onsite (Redmond, Washington) $152k - $241k/yr Full-time Mid Level 1 benefits + 1 perks
Posted 4 days ago

About the role

NVIDIA is seeking a Senior AI Engineer to design and optimize agentic AI systems for the CUDA ecosystem. This role involves collaborating with hardware and software teams to accelerate agent planning, tool-use, and code generation workloads using foundational models.

Skills

AI systems development CUDA C++ Python GPU programming Performance optimization Deep learning frameworks Inference stacks Agentic systems Orchestration frameworks Compiler integration Foundational models Multi-agent systems Software engineering Algorithm design
Onsite (San Jose, California) $250k - $286k/yr Full-time Senior Level 5 benefits + 3 perks
Posted 5 days ago

About the role

Capital One's Intelligent Foundations and Experiences team is building responsible, scalable AI systems to transform banking. This role involves designing and deploying GenAI platform services, including LLM training, inference, and optimization, to enhance customer and associate experiences.

Skills

Artificial Intelligence Machine Learning Python Go Scala Java Large Language Models Cloud Computing AWS VectorDBs PyTorch Model Evaluation System Optimization Software Engineering Leadership
Onsite (Jersey City, NJ) $204k - $285k/yr Full-time Senior Level 6 benefits + 2 perks
Posted 5 days ago

About the role

JPMorgan Chase is seeking a Principal Software Engineer to lead LLM inference optimization and benchmarking for its production AI platforms. This role focuses on driving performance, cost-efficiency, and scalability across the firm's enterprise AI infrastructure.

Skills

LLM inference GPU architecture Quantization Speculative decoding Benchmarking vLLM TensorRT-LLM SGLang AWS EKS Performance optimization Python Machine learning Distributed systems Chaos engineering Agentic AI
Onsite (San Jose, California) $176k - $265k/yr Full-time Senior Level 4 benefits + 3 perks
Posted 5 days ago

About the role

F5 is a global leader in application delivery and security, empowering organizations to create, secure, and run applications. The AI Inference Engineer role focuses on optimizing Large Language Models for high-performance inference, maximizing throughput and minimizing latency across diverse hardware environments.

Skills

Python C++ Rust Golang vLLM TensorRT Llama.cpp Ollama Docker Kubernetes AWS GCP Azure CUDA Triton MLOps
Onsite (Boston, Massachusetts) $167k - $260k/yr Full-time Senior Level 7 benefits + 2 perks
Posted 5 days ago

About the role

Join Amazon's AGI team to own the end-to-end inference stack for real-time multimodal conversational AI. You will co-design model architectures, optimize low-latency streaming runtimes, and build training infrastructure to make frontier-scale models viable at scale.

Skills

Inference optimization Deep learning Transformer architectures C++ Python GPU performance optimization CUDA Triton Multimodal AI Latency profiling Reinforcement learning Distributed training Quantization Model architecture Streaming systems TensorRT-LLM
Onsite (United States) $198k - $242k/yr Full-time Senior Level 4 benefits + 3 perks
Posted 5 days ago

About the role

Modular is building a next-generation AI platform to radically improve how developers build and deploy models. This role leads architecture decisions for distributed inference systems, designing robust, scalable abstractions that unify fragmented deployment processes.

Skills

Systems programming Performance tuning Concurrency Distributed architectures API design Python Rust Mojo Machine learning inference GPU programming Accelerator programming Dataflow programming Memory-layout optimization Defensive engineering System architecture
Onsite (Santa Clara, California) $184k - $356k/yr Full-time Senior Level 1 benefits + 1 perks
Posted 6 days ago

About the role

NVIDIA is seeking a Senior Software Engineer to optimize AI inference performance for LLMs and VLMs on GPU-accelerated systems. This role focuses on improving latency, throughput, and efficiency through kernel optimization and distributed runtime improvements.

Skills

LLM VLM CUDA Python Rust C++ GPU Architecture TensorRT-LLM vLLM Performance Profiling Distributed Systems Quantization Kernel Optimization Roofline Modeling Nsight Systems Nsight Compute
Onsite (San Jose, California) $164k - $313k/yr Full-time Senior Level 2 benefits + 1 perks
Posted 6 days ago

About the role

Adobe's Applied Science & Machine Learning team is seeking a Senior Applied Scientist/Engineer to bridge the gap between research and production. You will own the end-to-end training-to-deployment pipeline for next-generation video and multimodal generative models, ensuring they are reliable, performant, and cost-efficient in production.

Skills

Python PyTorch Distributed Training FSDP Tensor Parallelism Pipeline Parallelism Inference Optimization Generative Models Machine Learning GPU Efficiency Model Deployment Performance Profiling TensorRT vLLM Multimodal Models
Onsite (San Francisco, California) $266k - $445k/yr Full-time Senior Level 9 benefits + 6 perks
Posted 1 week ago

About the role

OpenAI is developing AI-native silicon and system-level solutions to power frontier models. This role involves building the low-level device runtime for custom AI accelerators, managing kernel scheduling, memory, and synchronization to enable efficient hardware execution.

Skills

C C++ Rust Systems programming Device drivers Firmware Concurrency Synchronization Memory management Accelerator software Cycle-accurate simulation Performance tuning Hardware-software co-design Debugging Kernel scheduling Data movement
Onsite (Denver, Colorado) $185k - $225k/yr Full-time Senior Level 10 benefits + 6 perks
Posted 1 week ago

About the role

Crusoe is an AI infrastructure company building vertically integrated solutions from energy to tokens. This role involves owning the end-to-end inference stack, optimizing large language models for performance and cost, and collaborating with customers to transition workloads to production.

Skills

Python C++ LLM inference vLLM SGLang CUDA Performance profiling GPU architecture Machine learning pipelines Distributed systems Docker Kubernetes Latency optimization Throughput optimization Production engineering
Hippocratic AI

LLM Inference Engineer

Hippocratic AI

Onsite (Menlo Park, California) Full-time Senior Level 2 benefits
Posted 1 week ago

About the role

Hippocratic AI is building a safety-focused large language model for the healthcare industry. This role involves owning and optimizing the serving infrastructure for LLMs to ensure high performance, low latency, and cost-efficiency for patient conversations.

Skills

LLM inference Distributed systems Python C++ CUDA Quantization Speculative decoding vLLM SGLang TensorRT-LLM Performance optimization GPU optimization Multi-LoRA serving Transformer models Infrastructure engineering
Remote (McLean, Virginia) $244k - $335k/yr Full-time Senior Level 5 benefits + 3 perks
Posted 1 week ago

About the role

Capital One's Intelligent Foundations and Experiences team is building responsible, scalable AI systems to transform banking. This role involves designing and deploying foundation models and LLM inference to enhance customer experiences and internal tools.

Skills

Artificial Intelligence Machine Learning Python Go Scala Java Large Language Models Cloud Computing AWS VectorDBs PyTorch Software Engineering System Architecture Model Evaluation Data Governance Leadership
Bose Corporation

Edge AI ML Engineer

Bose Corporation

Onsite (Framingham, Massachusetts) $141k - $193k/yr Full-time Senior Level 2 benefits + 2 perks
Posted 1 week ago

About the role

Bose Corporation is seeking an Edge AI ML Engineer to develop and optimize machine learning models for audio and multimodal intelligence on embedded devices. This role focuses on deploying efficient algorithms that meet real-world constraints like latency, memory, and power.

Skills

C++ C Machine learning Embedded systems Python Digital signal processing Quantization RTOS Firmware Neural networks Edge AI TFLite Micro ExecuTorch Audio processing Algorithm design Performance optimization
Onsite (San Francisco, California) Full-time Senior Level 2 benefits + 3 perks
Posted 1 week ago

About the role

Adaption is building efficient, evolving AI systems that adapt in real-time. This role involves owning the cost and performance of the inference stack by optimizing serving engines and managing core performance levers to improve throughput and latency.

Skills

Inference infrastructure Performance engineering Model serving KV-cache management Continuous batching Speculative decoding Quantization vLLM SGLang TensorRT-LLM Python C++ Rust CUDA NCCL GPU performance
Onsite (Mountain View, CA) $207k - $300k/yr Full-time Senior Level 1 benefits + 2 perks
Posted 1 week ago

About the role

Join Google DeepMind to push the boundaries of AI model execution at scale. You will analyze and optimize inference workloads to maximize hardware throughput, reduce latency, and drive systemic improvements across the fleet.

Skills

Python C++ AI model execution Inference optimization Distributed systems Profiling Benchmarking GPU performance TPU performance Quantization Collective communication Latency analysis Throughput optimization Memory bandwidth Observability Reliability
Onsite (San Jose, California) $229k - $286k/yr Full-time Senior Level 3 benefits + 2 perks
Posted 1 week ago

About the role

Capital One's Intelligent Foundations and Experiences team is building responsible, scalable AI systems to transform banking. This role involves designing and deploying foundational AI components, including LLM training and inference, to enhance customer experiences and internal tools.

Skills

Python Go Scala Java Machine learning Large language models Distributed systems Cloud platforms Vector databases PyTorch Huggingface AI infrastructure Model evaluation System optimization Leadership Communication
Onsite (Kirkland, WA) $147k - $210k/yr Full-time Mid Level 1 benefits + 2 perks
Posted 1 week ago

About the role

Join Google's Core team to build the technical foundation behind flagship products. This role involves developing next-generation technologies and infrastructure that impact billions of users worldwide.

Skills

Software Development Generative AI Large Language Models Multi-modal Models Large Vision Models Performance Debugging System Health Software Test Engineering Compiler Design Code Optimization High-performance Software Development Concurrent Programming Multi-core Computer Architectures Distributed Computing Artificial Intelligence Natural Language Processing
Onsite (San Jose, California) $250k - $286k/yr Full-time Senior Level 5 benefits + 3 perks
Posted 1 week ago

About the role

Capital One is seeking a Sr. Lead AI Engineer to design and deploy scalable AI software components, including foundation model training and LLM inference, for the Intelligent Foundations and Experiences team. This role drives the technical vision for foundational AI systems to enhance customer experiences and banking products.

Skills

Python Machine Learning Artificial Intelligence Large Language Models Cloud Computing AWS PyTorch VectorDBs Software Engineering System Optimization Leadership Mentoring Data Science Algorithm Development Communication Problem Solving

Finding more jobs