Software Engineer, Hardware Health

Slope

San Francisco (CA)

On-site

USD 130,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Slope is seeking an experienced engineer to maintain system integrity for its supercomputers used in AI research. The ideal candidate will have over 7 years in software engineering, proficiency in Python and shell scripting, and the ability to analyze complex data. Responsibilities include owning system health checks, diagnosing hardware failures, and developing automation tools to minimize disruptions during large-scale operations. This role is crucial for keeping advanced model training seamless and efficient.

Qualifications

  • 7+ years of industry experience in software engineering.
  • Proficiency with Python and shell scripting.
  • Experience developing reproducible analyses.

Responsibilities

  • Own and improve system health checks for supercomputers.
  • Lead investigations into hardware failures and bugs.
  • Build automation to monitor and fix issues.

Skills

Software engineering experience
Proficiency with Python
Shell scripting
Comfort with SQL
Data analysis with Pandas

Job description

About the Team

The Frontier Systems team at OpenAI builds, launches, and supports the largest supercomputers in the world that OpenAI uses for its most cutting edge model training.

We take data center designs, turn them into real, working systems and build any software needed for running large-scale frontier model trainings.

Our mission is to bring up, stabilize and keep these hyperscale supercomputers reliable and efficient during the training of the frontier models.

About the Role

On the Frontier Systems team, you’ll build critical infrastructure that keeps our supercomputers running reliably for cutting-edge AI research. Even a single hardware failure can derail a large-scale training run, so minimizing disruptions is core to the mission.

Engineers here own their work end-to-end and are trusted to make a real impact. This role is for someone who goes deep - who thrives on root-causing system-level issues and building automation to catch and fix problems at scale.

In this role, you will:
  • Own and improve the system health checks that keep our hyperscale supercomputers stable during model training.
  • Lead deep dives into hardware failures and system-level bugs to understand how things break at scale.
  • Build automation that monitors and fixes issues across thousands of machines - so researchers can keep moving without interruption.
You might thrive in this role if you have:
  • 7+ years of industry experience in software engineering
  • Proficiency with Python and shell scripting
  • A high degree of comfort digging into noisy data with SQL, PromQL, and Pandas or any other tool necessary
  • Experience developing reproducible analyses
  • A balance of strengths in building and operationalizing
Bonus if you have:
  • Experience with low level details of hardware components, protocols, and associated Linux tooling (e.g., PCIe, Infiniband, networking, power management, kernel perf tuning)
  • Experience with visualization of large data centers and networks.
  • Expertise with network operations and tooling
  • Expertise with power management and stabilization
Equal Opportunity Employer

We are an equal opportunity employer and do not discriminate on the basis of race, religion, national origin, gender, sexual orientation, age, veteran status, disability or any other legally protected status.

For US Based Candidates: Pursuant to the San Francisco Fair Chance Ordinance, we will consider qualified applicants with arrest and conviction records.

We are committed to providing reasonable accommodations to applicants with disabilities.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineer, Frontier Systems - Power Management
Software Engineer, Frontier Systems - Power Management

OpenAI • San Francisco (CA)

On-site
USD 295,000 - 445,000
Software Engineer, Hardware Health
Software Engineer, Hardware Health

OpenAI • San Francisco (CA)

On-site
USD 250,000 - 445,000
Software Engineer, Frontier Systems - Power Management
Software Engineer, Frontier Systems - Power Management

Slope • San Francisco (CA)

On-site
USD 310,000 - 460,000
Software Engineer, Compute Foundations Systems
Software Engineer, Compute Foundations Systems

OpenAI • San Francisco (CA)

On-site
USD 150,000 - 190,000
Software Engineer, Hardware
Software Engineer, Hardware

Slope • San Francisco (CA)

On-site
USD 310,000 - 460,000
Software Engineer, GPU Infrastructure - HPC
Software Engineer, GPU Infrastructure - HPC

OpenAI • San Francisco (CA)

On-site
USD 325,000 - 590,000
Software Engineer, Platform Systems
Software Engineer, Platform Systems

OpenAI • San Francisco (CA)

On-site
USD 310,000 - 460,000
Software Engineer, GPU Infrastructure - HPC
Software Engineer, GPU Infrastructure - HPC

OpenAI • New York (NY)

On-site
USD 150,000 - 210,000
Systems Software Engineer, Silicon Bringup
Systems Software Engineer, Silicon Bringup

AI Chopping Block • San Francisco (CA), Northern (KY)

Hybrid
USD 230,000 - 300,000
Hardware Operations Engineer
Hardware Operations Engineer

OpenAI • United States

On-site
USD 86,000 - 228,000