Software Engineer, GPU Infrastructure - HPC

Cloudjobs

New York (NY)

On-site

USD 120,000 - 160,000

Full time

5 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Cloudjobs in New York seeks a reliability engineer to ensure uptime and reliability of a large-scale compute fleet. You will build automation for provisioning, health monitoring, and incident response to minimize outages.

The role requires deep system knowledge, strong Linux and networking fundamentals, experience with server hardware, and proficiency in Python or Go. Collaboration with cross-functional teams is essential to maintain highly available infrastructure.

Qualifications

  • Experience managing large-scale server environments.
  • Proficiency in Python or Go.
  • Strong knowledge of Linux, networking and server hardware.
  • Ability to deep-dive into system-level investigations.

Responsibilities

  • Build automation for provisioning and scaling of servers.
  • Monitor health and performance of compute fleet.
  • Identify and fix bottlenecks to improve uptime.

Skills

Python
Go
Linux
SQL
PromQL
Pandas
HPC
Distributed Systems
Server Hardware
Networking
Automation
Prometheus
Grafana
IPMI
Redfish
PCIe

Job description

The role focuses on ensuring the reliability and uptime of OpenAI's compute fleet by minimizing hardware failures. Responsibilities include building automation for server provisioning, monitoring health, and fixing performance bottlenecks.


Requirements: Candidates should have experience managing large-scale server environments and proficiency in Python or Go. Strong knowledge of Linux, networking, and server hardware is required, with a capacity for deep system-level investigation.


Key Skills: Python, Go, Linux, SQL, PromQL, Pandas, HPC, Distributed Systems, Server Hardware, Networking, Automation, Prometheus, Grafana, IPMI, Redfish, PCIe

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineer, Fleet Hardware Health
Software Engineer, Fleet Hardware Health

Cloudjobs • San Francisco (CA)

On-site
USD 140,000 - 180,000
Software Engineer, GPU Infrastructure - HPC
Software Engineer, GPU Infrastructure - HPC

OpenAI • San Francisco (CA)

On-site
USD 325,000 - 590,000
Software Engineer, GPU Infrastructure - HPC
Software Engineer, GPU Infrastructure - HPC

OpenAI • New York (NY)

On-site
USD 150,000 - 210,000
GPU HPC Infrastructure Engineer | Scale & Automation
GPU HPC Infrastructure Engineer | Scale & Automation

OpenAI • New York (NY)

On-site
USD 150,000 - 210,000
HPC Infrastructure Engineer
HPC Infrastructure Engineer

Arcadia • San Francisco (CA)

On-site
USD 180,000 - 260,000
Sr HPC Hardware Engineer
Sr HPC Hardware Engineer

Career Techniques • Dallas (TX)

Hybrid
USD 120,000 - 180,000
Senior AI Infrastructure Engineer — Scale GPU Clusters Remote
Senior AI Infrastructure Engineer — Scale GPU Clusters Remote

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 280,000 - 420,000
Equity options
Health, vision, dental benefits
Unlimited PTO
+2
Software Engineer, AI Infra
Software Engineer, AI Infra

Makers Fund • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Equity
Health benefits
Monthly stipends
+1
Senior AI Cloud SRE — HPC & GPU Infrastructure
Senior AI Cloud SRE — HPC & GPU Infrastructure

Neura Market • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Platform Engineer
Platform Engineer

Harrison Clarke • San Francisco (CA)

On-site
USD 120,000 - 160,000