Software Engineer, Frontier Systems - Power Management

OpenAI

San Francisco (CA)

On-site

USD 295,000 - 445,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

OpenAI is seeking a Software Engineer focused on power management to support our large-scale supercomputing infrastructure. You will be tasked with developing solutions to optimize power usage, working closely with researchers and engineers.

The ideal candidate has over 7 years of experience with strong proficiency in Python and a background in electrical engineering concepts. This role offers a competitive compensation range of $295K to $445K.

Join us in enhancing the reliability and efficiency of cutting-edge research!

Qualifications

  • 7+ years of software engineering experience focused on large-scale challenges.
  • Proficiency in Python and familiarity with automation tools.
  • Experience with distributed systems and analyzing streaming data.

Responsibilities

  • Develop system-level solutions to optimize power usage in supercomputers.
  • Build automation to monitor power consumption during training workloads.
  • Design tools for real-time monitoring of power-related faults.

Skills

Python
Automation and scripting tools
Distributed systems
Electrical engineering concepts
Strong analytical skills

Tools

SQL
PromQL
Pandas

Job description

About The Team

The Frontier Systems team at OpenAI builds, launches, and supports the largest supercomputers in the world used for training cutting‑edge models. We design data centers, turn them into real systems, and build any software needed for large‑scale frontier model trainings. Our mission is to bring up, stabilize, and keep these hyperscale supercomputers reliable and efficient during training.

About The Role

As a Software Engineer focused on power management, you will work on critical infrastructure to support cutting‑edge research. With large‑scale supercomputers consuming substantial amounts of power, managing this efficiently is key to maximizing computational capacity. The role ensures that the supercomputing infrastructure runs smoothly while maintaining reliability and grid‑level power stability.

In This Role, You Will
  • Develop and implement system‑level and software‑level solutions to optimize power usage in large‑scale supercomputers, ensuring efficient and reliable operations.
  • Build automation to monitor power consumption patterns during training workloads and design algorithms to stabilize these fluctuations, preventing issues with grid reliability.
  • Work with researchers and engineers to design tools for real‑time monitoring, detection, and remediation of power‑related hardware and system faults.
  • Collaborate cross‑functionally to translate complex electrical system requirements into code, while driving continuous improvements in power management solutions.
  • Drive the development of power throttling mechanisms at the IT system level to dynamically adjust power usage based on workload demands and infrastructure limitations.
  • Collaborate with hardware design teams to integrate system‑level power control requirements into IT hardware design, ensuring seamless coordination between software‑driven power management and hardware capabilities.
You Might Thrive In This Role If You Have
  • 7+ years of software engineering experience with a focus on solving large‑scale, system‑level challenges.
  • Strong proficiency in Python and familiarity with automation and scripting tools (e.g., shell scripting).
  • Experience with distributed systems to efficiently aggregate and analyze streaming data.
  • Knowledge of electrical engineering concepts including digital signal processing, power systems, Fast Fourier Transforms, or related areas.
  • Experience in system‑level investigations and development of automated solutions to address power management, fault detection, and remediation.
  • Strong analytical skills and the ability to dig into noisy data (experience with SQL, PromQL, Pandas, etc.).
  • Comfort working with both hardware and software teams to solve multidisciplinary problems.
Bonus Points If You Have
  • Deep expertise with the power characteristics of synchronous workloads (as seen in supercomputing or model training environments).
  • Knowledge of power control requirements in IT hardware design, with the ability to drive cross‑functional collaboration to integrate power management features into hardware systems effectively.
  • Working knowledge of control system fundamentals and how physical systems respond to control strategies.
Equal Employment Opportunity Statement

OpenAI is an equal opportunity employer, and we do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, age, veteran status, disability, genetic information, or other applicable legally protected characteristic. For additional information, please see OpenAI’s affirmative action and equal employment opportunity policy statement. We are committed to providing reasonable accommodations to applicants with disabilities.

Compensation Range

$295K - $445K

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineer, Frontier Systems - Power Management
Software Engineer, Frontier Systems - Power Management

Slope • San Francisco (CA)

On-site
USD 310,000 - 460,000
Software Engineer, Hardware Health
Software Engineer, Hardware Health

Slope • San Francisco (CA)

On-site
USD 130,000 - 160,000
Power Management Engineer for Frontier Supercomputers
Power Management Engineer for Frontier Supercomputers

OpenAI • San Francisco (CA)

On-site
USD 295,000 - 445,000
Software Engineer, Hardware
Software Engineer, Hardware

Slope • San Francisco (CA)

On-site
USD 310,000 - 460,000
Software Engineer, GPU Infrastructure - HPC
Software Engineer, GPU Infrastructure - HPC

OpenAI • San Francisco (CA)

On-site
USD 325,000 - 590,000
Data Center Infrastructure Electrical Engineer
Data Center Infrastructure Electrical Engineer

OpenAI • Los Angeles (CA)

On-site
USD 257,000 - 327,000
Data Center Infrastructure Electrical Engineer
Data Center Infrastructure Electrical Engineer

OpenAI • Seattle (WA)

On-site
USD 257,000 - 327,000
System Power Engineer, Consumer Devices
System Power Engineer, Consumer Devices

jobr.pro • San Francisco (CA)

Hybrid
USD 120,000 - 150,000
Full Stack Engineer, Fleet Scheduling
Full Stack Engineer, Fleet Scheduling

OpenAI • San Francisco (CA)

Hybrid
USD 230,000 - 490,000
Software Engineer, Supercomputing
Software Engineer, Supercomputing

Thinkingmachines • San Francisco (CA)

On-site
USD 350,000 - 475,000
Generous health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1