Skip to content
AI Engineer Jobs
adaption

Inference Engineer

adaption

Location
Onsite (San Francisco, California)
Employment
Full-time
Level
Senior Level
Posted 3 days ago

About the Role

Adaption is building efficient, evolving AI systems that adapt in real-time. As an Inference Engineer, you will own the cost and performance of the inference stack, optimizing throughput, latency, and resource efficiency for model serving.

Skills

ML systems Inference infrastructure Performance engineering Model serving KV-cache management Continuous batching Speculative decoding Quantization vLLM SGLang TensorRT-LLM Python C++ Rust CUDA NCCL

Benefits

  • Medical benefits
  • Paid time off

Perks

  • Flexible work
  • Travel stipend
  • Lunch stipend

Full job details