Skip to content
AI Engineer Jobs
Nebius

Senior Machine Learning Engineer, LLM Inference Optimization

Nebius

Location
Onsite (Zurich, Zurich)
Employment
Full-time
Level
Senior Level
Posted 2 weeks ago

About the Role

Nebius is building a full-stack AI cloud platform to support developers and enterprises in training and deploying models. As a Senior Machine Learning Engineer, you will optimize LLM and VLM inference endpoints for latency, throughput, and cost efficiency in production.

Skills

Python PyTorch LLM Inference VLM vLLM SGLang TensorRT-LLM Triton Inference Server Quantization Transformer architecture GPU optimization Latency optimization Throughput optimization Memory efficiency CUDA Triton

Full job details