Skip to content
Skip to content
AI Engineer Jobs
Harell Data

Software Engineer, AI Infrastructure

Harell Data

Location
Onsite (Palo Alto, California)
Employment
Full-time
Level
Senior Level
Posted 6 days ago

About the Role

Harell Data is building intelligent infrastructure for scientific discovery, enabling secure sharing of proprietary datasets and high-performance compute. As an early engineer, you will own the GPU compute layer, inference systems, and ML pipelines, shaping the platform's technical direction while working directly with customers.

Skills

Kubernetes GPU Infrastructure AWS GCP Inference Systems ML Pipelines System Design Observability Model Serving Resource Scheduling Multi-tenancy Cost Management Data Ingestion Fine-tuning Incident Response Infrastructure Architecture

Full job details

About the Role

You'll be an early engineer reporting directly to the CTO. You'll own the compute layer: the GPU clusters and the inference systems that run on them. You'll make the architectural decisions that define the platform. You'll also work directly with customers to understand what they actually need and turn that into infrastructure that works at scale.

What You Will Do

  • Build the GPU compute layer - Orchestration for GPU workloads on Kubernetes: resource allocation, scheduling, multi-tenancy, and cost management.
  • Build the inference layer - Model loading, autoscaling, batching, and serving. You own the latency and throughput customers feel.
  • Own the ML pipeline end to end - Data ingestion, preprocessing, training and fine-tuning jobs, and recovery when multi-node jobs fail.
  • Work directly with customers - Debug fine-tuning jobs that fail or run slow. Build the observability that tracks model performance and resource health in real time.
  • Own reliability - Incident response, on-call, and keeping the platform up as usage grows.
  • Shape technical direction - Lead build-vs-buy decisions on infrastructure and security. Set engineering standards. Help hire the team you want to work with.

Qualifications

  • 5+ years building and operating production infrastructure, with a focus on ML workloads: training, inference, or data pipelines
  • Hands-on experience with Kubernetes on AWS or GCP, ideally with GPU workloads.
  • Strong CS fundamentals and system design chops
  • Comfortable with ambiguity — you've worked somewhere where the playbook didn't exist yet

 

Location note: this role is based in Palo Alto, CA. No relocation assistance available for this role.