Skip to content
AI Engineer Jobs
Baseten

Software Engineer- Inference Performance

Baseten

Location
Onsite (San Francisco, California)
Compensation
$180k - $360k/yr
Employment
Full-time
Level
Senior Level
Posted 3 days ago

About the Role

Baseten is an AI infrastructure platform enabling companies to ship AI products fast. This role involves optimizing LLM inference performance, improving GPU utilization, and building benchmarking frameworks to drive cost savings and efficiency.

Skills

Python C++ LLM optimization Quantization Speculative decoding Continuous batching PyTorch TensorRT TensorRT-LLM GPU architecture CUDA Triton CUTLASS Distributed serving Inference engines Performance profiling

Benefits

  • Medical insurance
  • Dental insurance
  • Vision insurance
  • 401(k)
  • Paid parental leave

Perks

  • Equity
  • Flexible PTO
  • Fertility stipend
  • Hybrid Work

Full job details