Senior GPU HPC Platform Reliability Engineer

OpenAI

San Francisco (CA)

On-site

USD 325,000 - 590,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A leading AI research company in San Francisco is seeking a software engineer for its Fleet High Performance Computing team. In this role, you'll ensure the reliability and uptime of the compute fleet, working with automation systems and monitoring tools. Ideal candidates have experience managing server environments and proficiency in languages like Python or Go. Join us to innovate in AI technology while maintaining high system efficiency.

Qualifications

  • Experience managing large-scale server environments.
  • A balance of strengths in building and operationalizing.
  • Strong Linux, networking, and server hardware knowledge.

Responsibilities

  • Build and maintain automation systems for provisioning and managing server fleets.
  • Develop tools to monitor server health and performance.
  • Identify and fix performance bottlenecks and inefficiencies.

Skills

Experience managing large-scale server environments
Proficiency in Python, Go, or similar languages
Strong Linux, networking, and server hardware knowledge
Comfort digging into noisy data with SQL, PromQL, and Pandas

Tools

Prometheus
Grafana

Job description

A leading AI research company in San Francisco is seeking a software engineer for its Fleet High Performance Computing team. In this role, you'll ensure the reliability and uptime of the compute fleet, working with automation systems and monitoring tools. Ideal candidates have experience managing server environments and proficiency in languages like Python or Go. Join us to innovate in AI technology while maintaining high system efficiency.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior GPU Fleet Reliability Engineer
Senior GPU Fleet Reliability Engineer

Fal • San Francisco (CA)

On-site
USD 180,000 - 250,000
Health, dental, and vision insurance
Learning and growth opportunities
Visa sponsorship and relocation assistance
+1
GPU HPC Fleet Reliability Engineer
GPU HPC Fleet Reliability Engineer

CoreWeave • Bellevue (WA)

Hybrid
USD 83,000 - 110,000
Senior HPC GPU Compute Engineer (Hybrid SF)
Senior HPC GPU Compute Engineer (Hybrid SF)

The San Francisco Compute Company • San Francisco (CA)

Hybrid
USD 180,000 - 260,000
Generous equity grant
Retirement matching
Comprehensive medical, dental, and vision insurance
+3
Fleet Reliability Engineer — HPC & GPU Clusters
Fleet Reliability Engineer — HPC & GPU Clusters

CoreWeave • Plano (TX)

Hybrid
USD 83,000 - 110,000
GPU Reliability Engineer - AI Supercomputing Fleet
GPU Reliability Engineer - AI Supercomputing Fleet

Thinkingmachines • San Francisco (CA)

On-site
USD 350,000 - 475,000
Unlimited PTO
Paid parental leave
Relocation support
+1
Senior GPU HPC Systems Engineer
Senior GPU HPC Systems Engineer

Harvey Nash • Chicago (IL)

Hybrid
USD 125,000 - 150,000
Software Engineer, Scalable AI HPC Platform
Software Engineer, Scalable AI HPC Platform

Magic • United States

On-site
USD 200,000 - 550,000
Remote HPC AI Solutions Architect for Research Compute
Remote HPC AI Solutions Architect for Research Compute

Cambridge Computer Services, Inc • San Francisco (CA)

Hybrid
USD 90,000 - 210,000
Competitive salary
Multiple health insurance options
401(k) savings plan with employer matching
+2
Senior AI Cloud SRE — HPC & GPU Infrastructure
Senior AI Cloud SRE — HPC & GPU Infrastructure

Neura Market • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Senior AI Infrastructure Performance Engineer
Senior AI Infrastructure Performance Engineer

Crusoe • San Francisco (CA)

On-site
USD 172,000 - 210,000