Skip to content
AI Engineer Jobs
Nebius

Senior Machine Learning Engineer, LLM Inference Optimization

Nebius

Location
Onsite (London, England)
Employment
Full-time
Level
Senior Level
Posted 2 weeks ago

About the Role

Nebius is building a full-stack AI cloud platform to support developers and enterprises in training and deploying models. This role focuses on optimizing LLM and VLM inference services to improve latency, throughput, and cost efficiency for frontier models.

Skills

Python PyTorch LLM Inference Optimization VLM Inference Inference Engines Transformer Inference Quantization Distillation Speculative Decoding KV-Cache Optimization Benchmarking Performance Optimization GPU Utilization CUDA Triton Technical Communication

Full job details