Skip to content
AI Engineer Jobs
Onsite (Clermont-Dessous, Nouvelle-Aquitaine) $292k - $507k/yr Full-time Senior Level 1 benefits + 1 perks
Posted 1 day ago

About the role

NVIDIA seeks a Senior Solutions Architect to guide EMEA customers in deploying and optimizing large-scale AI inference workloads on multi-node GPU clusters. This role involves collaborating with product teams to improve inference efficiency and leading technical community engagement.

Skills

AI inference Neural networks Transformer optimization Quantization Speculative decoding KV cache management MoE inference GPU clusters High-performance computing TensorRT-LLM NVLink InfiniBand RDMA UCX Distributed systems Load balancing
Onsite (Mountain View, CA) $207k - $300k/yr Full-time Senior Level 1 benefits + 2 perks
Posted 1 day ago

About the role

Join Google's AI Rapid Response Team as a Staff AI/ML Engineer to lead technical architecture and efficiency strategies for high-leverage machine learning workloads. You will serve as a primary technical anchor, translating complex executive mandates into rigorous engineering solutions through rapid prototyping and embedded engagements.

Skills

Software development Machine learning C++ Python System design Distributed computing Model compression Inference optimization XLA CUDA Triton Tensor processing units Generative AI LLM interfaces Technical leadership Data structures
Hybrid (London, England) Full-time Senior Level 1 perks
Posted 1 day ago

About the role

Isomorphic Labs is applying frontier AI to accelerate drug discovery and cure diseases. This role involves implementing and optimizing large-scale LLM post-training methods and distributed training systems to support scientific breakthroughs in digital biology.

Skills

Large scale distributed training LLM JAX PyTorch NCCL GPU architectures Performance optimization Supervised fine-tuning Reinforcement learning Distributed training Inference systems Low-precision methods XLA Triton Pallas CUDA
Hybrid (Israel) Full-time Senior Level 1 perks
Posted 2 days ago

About the role

Join Microsoft's AI Frameworks Group to build the software foundation enabling efficient AI model deployment across hardware platforms. This high-impact role focuses on optimizing performance, scalability, and efficiency for Microsoft's AI workloads through advanced compilation and runtime innovations.

Skills

C++ Python AI Optimization AI Kernels AI Compilation Machine Learning Distributed Execution Quantization Sparsity Speculative Decoding Graph Optimization Operator Fusion Code Generation Inference Serving Batching Scheduling
Hybrid (Israel) Full-time Senior Level 1 perks
Posted 2 days ago

About the role

Microsoft's AI Frameworks Group is seeking a Principal AI Optimization Software Engineer Lead to build the software foundation for efficient AI model deployment across hardware platforms. This high-impact role focuses on optimizing performance, scalability, and efficiency for Microsoft's AI workloads through advanced compilation and hardware optimization.

Skills

C++ Python AI Optimization Machine Learning Compiler Design Distributed Systems Hardware Acceleration Quantization Sparsity Speculative Decoding Graph Optimization Operator Fusion Inference Serving Software Architecture Performance Engineering Technical Leadership
Onsite (Shanghai, Shanghai,CN) Full-time Senior Level
Posted 2 days ago

About the role

Qualcomm is seeking an AI SDK Software Engineer to develop and optimize neural network operator kernels and SDK features for automotive hardware platforms. The role focuses on integrating high-performance software with advanced hardware to enhance ADAS and GenAI applications using Snapdragon chipsets.

Skills

C++ Python Deep Learning Neural Networks PyTorch ONNX Hexagon DSP SIMD Quantization Inference Engine Embedded Systems ARM Architecture ADAS GenAI Software Development Performance Optimization
Onsite (Guangdong, Guangdong Province,CN) Full-time Mid Level
Posted 2 days ago

About the role

Qualcomm is seeking a Voice AI Software Engineer to design and optimize voice technologies, including ASR and TTS, for Snapdragon platforms. The role focuses on delivering low-power, high-performance on-device AI solutions that serve as front-end components for large language models.

Skills

Android System Development Python C++ Java Machine Learning TensorFlow PyTorch Keras Model Optimization Quantization Pruning Voice AI ASR TTS Snapdragon Platforms
Firmus Technologies

AI Engineer - Inference

Firmus Technologies

Onsite (Sydney, New South Wales) Full-time Senior Level
Posted 2 days ago

About the role

Firmus Technologies is building efficient, sustainable AI infrastructure through its AI Factory and AI Cloud platform. This role involves building and optimizing self-hosted AI inference services to provide scalable, high-performance endpoints for internal and external applications.

Skills

AI Inference Model Serving TensorRT-LLM vLLM Python C++ Go Kubernetes CUDA GPU Profiling Distributed Systems Quantization Performance Benchmarking CI/CD LLM Infrastructure Optimization
SatoshiLabs

AI Platform Engineer

SatoshiLabs

Onsite (Prague, Prague) Full-time Mid Level 6 perks
Posted 2 days ago

About the role

SatoshiLabs, the creator of Trezor hardware wallets, is building an internal AI inference infrastructure to allow employees to use large language models securely on-premises. The AI Platform Engineer will own the end-to-end lifecycle of this GPU-based system, managing hardware, OS, and model serving stacks while supporting developers.

Skills

Ansible Terraform Docker Kubernetes Linux Networking Storage Containers Systemd Troubleshooting LLM vLLM SGLang NVIDIA GPU CUDA Infrastructure as Code
Majestic Labs

AI Kernel Engineer

Majestic Labs

Onsite (Raanana, Israel) Full-time Senior Level
Posted 2 days ago

About the role

Majestic Labs is a US-Israeli AI startup building next-generation infrastructure for large-scale AI workloads. The role involves designing and optimizing high-performance compute kernels for AI primitives to accelerate intelligence.

Skills

C++ CUDA Triton Parallel programming SIMD Performance profiling Kernel optimization Memory hierarchy Matrix engines RISC-V PyTorch LLVM MLIR Numerical stability Linear algebra Debugging
Onsite (Boston, Massachusetts) Full-time Mid Level 7 benefits + 5 perks
Posted 5 days ago

About the role

Netpreme is an early-stage AI infrastructure company building efficient memory processing units for large language models. This role involves deploying and optimizing LLM serving infrastructure using vLLM and SGLang on Kubernetes, focusing on performance engineering and system architecture to unlock greater AI capability.

Skills

vLLM SGLang Kubernetes Python LLM inference GPU systems Performance engineering Distributed GPU execution Continuous batching KV cache Quantization Speculative decoding CUDA Graphs Profiling Parallelism strategies
Onsite (San Jose, California) $250k - $286k/yr Full-time Senior Level 5 benefits + 3 perks
Posted 6 days ago

About the role

Capital One's Intelligent Foundations and Experiences team is building responsible, scalable AI systems to transform banking. This role involves designing and deploying foundation model training, LLM inference, and guardrails to enhance customer experiences.

Skills

Artificial Intelligence Machine Learning Python Go Scala Java LLM Inference VectorDBs Cloud Computing AWS PyTorch Nemo Guardrails Software Engineering System Optimization Leadership Communication
Hybrid (Tustin, California) $95k - $113k/yr Full-time Senior Level 8 benefits + 4 perks
Posted 6 days ago

About the role

Advantech is a global leader in IoT intelligent systems and embedded platforms. This role focuses on developing and deploying robotics and Edge AI software solutions, optimizing AI vision algorithms for edge deployment, and integrating sensor data for autonomous navigation.

Skills

Robotics ROS2 Computer vision Embedded Linux AI inference Python C++ C Docker LiDAR SLAM TensorRT ONNX Runtime OpenVINO Gazebo Navigation
Onsite (San Jose, California) $200k - $450k/yr Full-time Senior Level
Posted 1 week ago

About the role

Hark is building advanced, personalized AI intelligence and next-generation hardware to create a universal interface between humans and machines. This role focuses on optimizing transformer models to run efficiently on constrained embedded hardware.

Skills

C++ C SIMD Kernel Optimization Transformer Models Inference Optimization DSP NPU GPU Memory Management Profiling Latency Optimization Power Budgeting Embedded Systems Compiler Toolchains
Onsite (San Jose, California) $229k - $286k/yr Full-time Senior Level 5 benefits + 3 perks
Posted 1 week ago

About the role

Capital One's Intelligent Foundations and Experiences team is building responsible, scalable AI systems to transform banking. This role involves designing and deploying foundational AI components, including LLM inference and model evaluation, to enhance customer and associate experiences.

Skills

Python Go Scala Java Machine learning LLM inference VectorDBs PyTorch AWS Cloud computing Model evaluation Guardrails Software engineering Mathematics System optimization Leadership
Onsite (Jersey City, NJ) $204k - $285k/yr Full-time Senior Level 6 benefits + 1 perks
Posted 1 week ago

About the role

JPMorgan Chase seeks a Principal Software Engineer to lead the design and evolution of its GenAI serving platform. This role focuses on high-performance LLM inference, GPU efficiency, and intelligent model routing to deliver reliable, cost-effective AI capabilities at enterprise scale.

Skills

LLM inference Python Java Scala Go GPU optimization Distributed systems Cloud-native Model routing Performance engineering Quantization API design System architecture Observability Agile methodologies Infrastructure as Code
Capital One
Onsite (San Jose, California) $343k - $392k/yr Full-time Senior Level 5 benefits + 3 perks
Posted 1 week ago

About the role

Capital One's Intelligent Foundations and Experiences team is building responsible, scalable AI systems to transform banking. This role involves designing and deploying foundation models, LLM inference, and guardrails to enhance customer and associate experiences.

Skills

Machine learning Artificial intelligence Python Go Scala Java Large language models Cloud computing AWS PyTorch VectorDBs Nemo guardrails Huggingface System architecture Software engineering Mathematical modeling
Onsite (Chicago, Illinois) $190k - $225k/yr Full-time Senior Level
Posted 1 week ago

About the role

AssemblyAI is a capital-efficient Voice AI company processing millions of daily API calls, seeking a Senior Software Engineer to productionize research models into scalable, customer-facing APIs. The role involves bridging research and product teams to build robust infrastructure and SDKs for a fast-paced, meritocratic team.

Skills

Backend Engineering Machine Learning Infrastructure API Development System Scalability Production Systems Software Architecture Collaboration Technical Documentation Inference Optimization Product Design SDK Development Internal Tooling
Onsite (Seattle, WA) $175k - $260k/yr Full-time Senior Level 3 benefits + 3 perks
Posted 1 week ago

About the role

JPMorgan Chase seeks a Senior Lead Software Engineer to architect and deploy secure, scalable cloud platforms optimized for AI/ML workloads. This role drives business impact by collaborating with cross-functional teams to enhance technology products and operational efficiency.

Skills

Kubernetes Docker Python Go Java C# Cloud computing Infrastructure as code Microservices Machine learning Transformer architecture CI/CD System design GPU infrastructure Observability MLOps
Onsite (Culver City, California) $143k - $213k/yr Full-time Mid Level 7 benefits + 2 perks
Posted 1 week ago

About the role

Join the AI Studios Engineering org within Prime Video and Amazon MGM Studios to build and operate ML training and serving infrastructure for the CreativeFlux platform. You will deploy scalable inference pipelines and optimize GPU utilization to support generative models in production for professional animation and VFX.

Skills

Machine learning Kubernetes EKS Generative AI Python C++ Java Distributed systems GPU optimization Inference pipelines SageMaker Model training Model serving Autoscaling Infrastructure as code Software engineering

Finding more jobs