Software Engineer, AI Infrastructure

Harell Data

Palo Alto (CA)

On-site

USD 180,000 - 260,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Harell Data is seeking an experienced engineer to lead the compute platform in Palo Alto, focusing on GPU clusters and scalable ML infrastructure. You will own the end-to-end ML pipeline, from data ingestion to deployment, while guiding architectural choices and collaborating closely with customers to translate needs into reliable systems.

You will pair with the CTO and shape the engineering direction, implementing Kubernetes-based orchestration, autoscaling, and efficient resource management to

Qualifications

  • 5+ years building and operating production infrastructure, with a focus on ML workloads
  • Hands-on experience with Kubernetes on AWS or GCP, ideally with GPU workloads
  • Strong CS fundamentals and system design chops
  • Comfortable with ambiguity — you’ve worked somewhere where the playbook didn’t exist yet

Responsibilities

  • Build the GPU compute layer — orchestration for GPU workloads on Kubernetes, including resource allocation and cost management
  • Build the inference layer — model loading, autoscaling, batching, and serving with low latency
  • Own the ML pipeline end-to-end — data ingestion, preprocessing, training, fine-tuning, and recovery
  • Work directly with customers to debug fine-tuning jobs and improve observability of model performance and resource health
  • Own reliability through incident response, on-call duties, and platform uptime as usage grows
  • Shape technical direction by leading build-vs-buy decisions and setting engineering standards

Skills

Kubernetes
ML infrastructure
System design
Ambiguity tolerance
GPU compute
AWS/GCP

Tools

GPU clusters
Kubernetes on AWS/GCP

Job description

You\'ll be an early engineer reporting directly to the CTO. You\'ll own the compute layer: the GPU clusters and the inference systems that run on them. You\'ll make the architectural decisions that define the platform. You\'ll also work directly with customers to understand what they actually need and turn that into infrastructure that works at scale.

What You Will Do

  • Build the GPU compute layer - Orchestration for GPU workloads on Kubernetes: resource allocation, scheduling, multi-tenancy, and cost management.
  • Build the inference layer - Model loading, autoscaling, batching, and serving. You own the latency and throughput customers feel.
  • Own the ML pipeline end to end - Data ingestion, preprocessing, training and fine-tuning jobs, and recovery when multi-node jobs fail.
  • Work directly with customers - Debug fine-tuning jobs that fail or run slow. Build the observability that tracks model performance and resource health in real time.
  • Own reliability - Incident response, on-call, and keeping the platform up as usage grows.
  • Shape technical direction - Lead build-vs-buy decisions on infrastructure and security. Set engineering standards. Help hire the team you want to work with.

Qualifications

  • 5+ years building and operating production infrastructure, with a focus on ML workloads: training, inference, or data pipelines
  • Hands-on experience with Kubernetes on AWS or GCP, ideally with GPU workloads.
  • Strong CS fundamentals and system design chops
  • Comfortable with ambiguity — you\'ve worked somewhere where the playbook didn\'t exist yet

Location note: this role is based in Palo Alto, CA. No relocation assistance available for this role.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Platform Engineer
Platform Engineer

Harrison Clarke • San Francisco (CA)

On-site
USD 120,000 - 160,000
Platform Engineer
Platform Engineer

Reactor • San Francisco (CA)

On-site
USD 190,000 - 260,000
Visa sponsorship
Relocation assistance
Health, dental, vision
+1
AI Infrastructure Engineer
AI Infrastructure Engineer

dicedemo • Boston (CT)

On-site
USD 130,000 - 170,000
Software Engineer, Compute Foundations
Software Engineer, Compute Foundations

Linuxcareers • San Francisco (CA), Northern (KY)

On-site
USD 210,000 - 270,000
Founding Engineer - ML Infrastructure
Founding Engineer - ML Infrastructure

uRun • San Francisco (CA)

On-site
USD 120,000 - 160,000
Health, dental, and vision
401(k)
Paid time off
+2
Member of Technical Staff, AI Compute & Data Infrastructure Vinci
Member of Technical Staff, AI Compute & Data Infrastructure Vinci

CDFAM - Computational Design Symposium • Palo Alto (CA), Northern (KY)

On-site
USD 180,000 - 240,000
Member of Technical Staff, Platform - AI Infrastructure
Member of Technical Staff, Platform - AI Infrastructure

Hamilton Barnes Associates Limited • United States

On-site
USD 213,000 - 288,000
Equity
Health care
Member of Technical Staff - GPU Infrastructure
Member of Technical Staff - GPU Infrastructure

Prime Intellect • San Francisco (CA), Northern (KY)

On-site
USD 150,000 - 300,000
Equity incentives
Senior Infrastructure Engineer
Senior Infrastructure Engineer

Harrison Clarke • United States

On-site
USD 100,000 - 140,000
AI Infrastructure Engineer
AI Infrastructure Engineer

AI Breaking Wire • Menlo Park (CA), Northern (KY)

On-site
USD 200,000 - 350,000
RSUs
Health benefits
Parental leave
+1