GPU Infra Engineer: Scale Massive Clusters & Observability

Exa

San Francisco (CA)

On-site

USD 120,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A cutting-edge tech company in San Francisco seeks infrastructure engineers to enhance the tooling and systems that power its AI applications. Responsibilities include building GPU orchestration, scaling cloud batchjob systems, and designing efficient scheduling software. Candidates should have experience with large-scale infrastructure and a strong focus on reliability and observability. This position is in-person, and international sponsorship is available.

Qualifications

  • Experience with large-scale infrastructure, such as GPU clusters or cloud batchjob systems.
  • Obsessive focus on reliability, observability, and optimization.

Responsibilities

  • Build GPU cluster orchestration.
  • Scale AWS batchjob systems to handle extensive jobs.
  • Design GPU scheduling software for maximum utilization.
  • Develop observability tooling for production systems.

Skills

Designing and operating large-scale infrastructure
Reliability and observability mindset
Optimization across the entire stack

Tools

Kubernetes
AWS batchjob systems
GPU scheduling software

Job description

A cutting-edge tech company in San Francisco seeks infrastructure engineers to enhance the tooling and systems that power its AI applications. Responsibilities include building GPU orchestration, scaling cloud batchjob systems, and designing efficient scheduling software. Candidates should have experience with large-scale infrastructure and a strong focus on reliability and observability. This position is in-person, and international sponsorship is available.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Infra & Cluster Engineer — Scale GPU/CPU Orchestration
AI Infra & Cluster Engineer — Scale GPU/CPU Orchestration

Linuxcareers • San Francisco (CA)

On-site
USD 120,000 - 160,000
GPU Infra Solutions Architect for Large-Scale AI Clusters
GPU Infra Solutions Architect for Large-Scale AI Clusters

Prime Intellect • San Francisco (CA)

On-site
USD 150,000 - 300,000
Senior GPU Cloud Infrastructure Engineer
Senior GPU Cloud Infrastructure Engineer

Hyperbolic • San Francisco (CA)

On-site
USD 180,000 - 260,000
Distributed Systems Engineer - AI Infra & GPU Clusters
Distributed Systems Engineer - AI Infra & GPU Clusters

krea.ai • San Francisco (CA)

On-site
USD 120,000 - 160,000
Staff GPU Infra Engineer: Terraform, K8s & Large-Scale AI
Staff GPU Infra Engineer: Terraform, K8s & Large-Scale AI

Magic AI, Inc • San Francisco (CA)

On-site
USD 200,000 - 550,000
Equity
401(k) with 6% match
Health, dental and vision insurance
+3
Senior AI Infra Engineer: GPU Clusters & Kubernetes
Senior AI Infra Engineer: GPU Clusters & Kubernetes

Intelliswift - An LTTS Company • Sunnyvale (CA)

On-site
USD 120,000 - 150,000
GPU Infra Engineer for Scalable AI Compute
GPU Infra Engineer for Scalable AI Compute

OpenAI • United States

On-site
USD 180,000 - 240,000
Senior Observability Platform Engineer – GPU/AI Infra
Senior Observability Platform Engineer – GPU/AI Infra

Nscale • San Francisco (CA)

On-site
USD 130,000 - 160,000
ML Infra Engineer: Scale GPU Training & Inference
ML Infra Engineer: Scale GPU Training & Inference

Reducto • San Francisco (CA)

On-site
USD 120,000 - 160,000
Unlimited PTO
Free lunch
Reimbursed transportation
+3
Lead Large-Scale GPU Cluster Engineer for AI Research
Lead Large-Scale GPU Cluster Engineer for AI Research

Linuxcareers • San Francisco (CA)

On-site
USD 120,000 - 180,000