Software Engineer, AI Infrastructure

Harell Data

Palo Alto (CA)

On-site

USD 180,000 - 260,000

Full time

13 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Harell Data is seeking an experienced engineer to lead the compute platform in Palo Alto, focusing on GPU clusters and scalable ML infrastructure. You will own the end-to-end ML pipeline, from data ingestion to deployment, while guiding architectural choices and collaborating closely with customers to translate needs into reliable systems.

You will pair with the CTO and shape the engineering direction, implementing Kubernetes-based orchestration, autoscaling, and efficient resource management to

Qualifications

  • 5+ years building and operating production infrastructure, with a focus on ML workloads
  • Hands-on experience with Kubernetes on AWS or GCP, ideally with GPU workloads
  • Strong CS fundamentals and system design chops
  • Comfortable with ambiguity — you’ve worked somewhere where the playbook didn’t exist yet

Responsibilities

  • Build the GPU compute layer — orchestration for GPU workloads on Kubernetes, including resource allocation and cost management
  • Build the inference layer — model loading, autoscaling, batching, and serving with low latency
  • Own the ML pipeline end-to-end — data ingestion, preprocessing, training, fine-tuning, and recovery
  • Work directly with customers to debug fine-tuning jobs and improve observability of model performance and resource health
  • Own reliability through incident response, on-call duties, and platform uptime as usage grows
  • Shape technical direction by leading build-vs-buy decisions and setting engineering standards

Skills

Kubernetes
ML infrastructure
System design
Ambiguity tolerance
GPU compute
AWS/GCP

Tools

GPU clusters
Kubernetes on AWS/GCP

Job description

You\'ll be an early engineer reporting directly to the CTO. You\'ll own the compute layer: the GPU clusters and the inference systems that run on them. You\'ll make the architectural decisions that define the platform. You\'ll also work directly with customers to understand what they actually need and turn that into infrastructure that works at scale.

What You Will Do

  • Build the GPU compute layer - Orchestration for GPU workloads on Kubernetes: resource allocation, scheduling, multi-tenancy, and cost management.
  • Build the inference layer - Model loading, autoscaling, batching, and serving. You own the latency and throughput customers feel.
  • Own the ML pipeline end to end - Data ingestion, preprocessing, training and fine-tuning jobs, and recovery when multi-node jobs fail.
  • Work directly with customers - Debug fine-tuning jobs that fail or run slow. Build the observability that tracks model performance and resource health in real time.
  • Own reliability - Incident response, on-call, and keeping the platform up as usage grows.
  • Shape technical direction - Lead build-vs-buy decisions on infrastructure and security. Set engineering standards. Help hire the team you want to work with.

Qualifications

  • 5+ years building and operating production infrastructure, with a focus on ML workloads: training, inference, or data pipelines
  • Hands-on experience with Kubernetes on AWS or GCP, ideally with GPU workloads.
  • Strong CS fundamentals and system design chops
  • Comfortable with ambiguity — you\'ve worked somewhere where the playbook didn\'t exist yet

Location note: this role is based in Palo Alto, CA. No relocation assistance available for this role.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Member of Technical Staff (Software Engineer, GPU Cluster Infrastructure)
Member of Technical Staff (Software Engineer, GPU Cluster Infrastructure)

United States Digital Space LLC • San Francisco (CA)

On-site
USD 180,000 - 260,000
Staff Software Engineer (AI Infrastructure)
Staff Software Engineer (AI Infrastructure)

DeepRec.ai • Palo Alto (CA)

On-site
USD 180,000 - 320,000
Platform Engineer
Platform Engineer

Harrison Clarke • San Francisco (CA)

On-site
USD 120,000 - 160,000
Member of Technical Staff (Software Engineer, Inference & Training Platform)
Member of Technical Staff (Software Engineer, Inference & Training Platform)

United States Digital Space LLC • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Member of Technical Staff (Software Engineer, GPU Cluster Infrastructure)
Member of Technical Staff (Software Engineer, GPU Cluster Infrastructure)

Perplexity • New York (NY)

On-site
USD 250,000 - 485,000
Member of Technical Staff (Software Engineer, GPU Cluster Infrastructure)
Member of Technical Staff (Software Engineer, GPU Cluster Infrastructure)

B Capital • United States

On-site
USD 180,000 - 230,000
Founding Engineer - ML Platforms
Founding Engineer - ML Platforms

Cumulus Labs (YC W26) • San Francisco (CA)

On-site
USD 180,000 - 260,000
Member of Technical Staff – Software Engineer, GPU Cluster Infrastructure
Member of Technical Staff – Software Engineer, GPU Cluster Infrastructure

Perplexity • San Francisco (CA)

On-site
USD 180,000 - 240,000
Sr. Platform Engineer, ML Infrastructure
Sr. Platform Engineer, ML Infrastructure

Insilico Search Partners • Cambridge (MA)

On-site
USD 140,000 - 210,000
AI/ML Infra Engineer - Hosting
AI/ML Infra Engineer - Hosting

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 225,000 - 275,000
Stock options