Staff Infra Engineer – Large-Scale GPU Training

Hark

San Jose (CA)

On-site

USD 180,000 - 450,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Hark in San Jose, California, is seeking a Member of Technical Staff, Infrastructure Compute to manage large-scale GPU computing clusters for AI training and deployment. This critical role requires extensive experience in infrastructure and systems engineering along with strong knowledge of machine learning environments.

The ideal candidate will have a proven track record in managing distributed compute infrastructure, focusing on reliability and scalability. Competitive compensation package is offered for this full-time role.

Qualifications

  • 5+ years of experience in infrastructure, systems, or platform engineering.
  • Experience managing GPU clusters or large-scale distributed compute infrastructure.
  • Strong proficiency in systems or infrastructure programming languages.

Responsibilities

  • Design and maintain Infrastructure as Code (IaC) for cluster provisioning.
  • Enhance CI/CD pipelines for model service delivery.
  • Own stable training infrastructure for 10,000+ GPUs.

Skills

Systems engineering
Machine learning infrastructure
GPU management
Container orchestration

Tools

Kubernetes
Pulumi

Job description

Hark in San Jose, California, is seeking a Member of Technical Staff, Infrastructure Compute to manage large-scale GPU computing clusters for AI training and deployment. This critical role requires extensive experience in infrastructure and systems engineering along with strong knowledge of machine learning environments.

The ideal candidate will have a proven track record in managing distributed compute infrastructure, focusing on reliability and scalability. Competitive compensation package is offered for this full-time role.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior ML Infra Architect — Large-Scale GPU Training
Senior ML Infra Architect — Large-Scale GPU Training

Hark, Inc. • San Jose (CA)

On-site
USD 180,000 - 450,000
Infrastructure, Large-scale Training San Jose
Infrastructure, Large-scale Training San Jose

Hark, Inc. • San Jose (CA)

On-site
USD 180,000 - 450,000
Senior Training Infra Engineer - 800+ GPU Scale
Senior Training Infra Engineer - 800+ GPU Scale

Figureai • San Jose (CA)

On-site
USD 150,000 - 350,000
GPU Systems Engineer for AI Training Clusters
GPU Systems Engineer for AI Training Clusters

Thinkingmachines • San Francisco (CA)

On-site
USD 350,000 - 475,000
Generous health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
Staff Compute Infra Engineer - GPU & AI Systems
Staff Compute Infra Engineer - GPU & AI Systems

xAI • Palo Alto (CA)

On-site
USD 180,000 - 440,000
Lead Large-Scale GPU Cluster Engineer for AI Research
Lead Large-Scale GPU Cluster Engineer for AI Research

Linuxcareers • San Francisco (CA)

On-site
USD 120,000 - 180,000
Infrastructure, Large-scale Training
Infrastructure, Large-scale Training

Hark • San Jose (CA)

On-site
USD 180,000 - 450,000
Staff GPU Infra Engineer: Terraform, K8s & Large-Scale AI
Staff GPU Infra Engineer: Terraform, K8s & Large-Scale AI

Magic AI, Inc • San Francisco (CA)

On-site
USD 200,000 - 550,000
Equity
401(k) with 6% match
Health, dental and vision insurance
+3
Remote Engineering Manager, AI GPU Infrastructure
Remote Engineering Manager, AI GPU Infrastructure

5C • United States

On-site
USD 180,000 - 220,000
GPU Infrastructure Engineer — Scalable AI Training
GPU Infrastructure Engineer — Scalable AI Training

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 350,000 - 475,000
Health benefits
Unlimited PTO
Parental leave
+1