GPU Infrastructure Reliability Engineer

Cloudjobs

New York (NY)

On-site

USD 120,000 - 160,000

Full time

5 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Cloudjobs in New York seeks a reliability engineer to ensure uptime and reliability of a large-scale compute fleet. You will build automation for provisioning, health monitoring, and incident response to minimize outages.

The role requires deep system knowledge, strong Linux and networking fundamentals, experience with server hardware, and proficiency in Python or Go. Collaboration with cross-functional teams is essential to maintain highly available infrastructure.

Qualifications

  • Experience managing large-scale server environments.
  • Proficiency in Python or Go.
  • Strong knowledge of Linux, networking and server hardware.
  • Ability to deep-dive into system-level investigations.

Responsibilities

  • Build automation for provisioning and scaling of servers.
  • Monitor health and performance of compute fleet.
  • Identify and fix bottlenecks to improve uptime.

Skills

Python
Go
Linux
SQL
PromQL
Pandas
HPC
Distributed Systems
Server Hardware
Networking
Automation
Prometheus
Grafana
IPMI
Redfish
PCIe

Job description

Cloudjobs in New York seeks a reliability engineer to ensure uptime and reliability of a large-scale compute fleet. You will build automation for provisioning, health monitoring, and incident response to minimize outages.

The role requires deep system knowledge, strong Linux and networking fundamentals, experience with server hardware, and proficiency in Python or Go. Collaboration with cross-functional teams is essential to maintain highly available infrastructure.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Software Engineer - GPU Fleet & Cloud Infra
Staff Software Engineer - GPU Fleet & Cloud Infra

Cloudjobs • San Francisco (CA)

On-site
USD 150,000 - 210,000
Restricted Stock Units
Health insurance
Vision insurance
+15
Software Engineer, GPU Infrastructure - HPC
Software Engineer, GPU Infrastructure - HPC

Cloudjobs • New York (NY)

On-site
USD 120,000 - 160,000
Software Engineer - Fleet Hardware Health & Automation
Software Engineer - Fleet Hardware Health & Automation

Cloudjobs • San Francisco (CA)

On-site
USD 140,000 - 180,000
GPU HPC Infrastructure Engineer | Scale & Automation
GPU HPC Infrastructure Engineer | Scale & Automation

OpenAI • New York (NY)

On-site
USD 150,000 - 210,000
HPC Fleet Reliability Engineer
HPC Fleet Reliability Engineer

CoreWeave • Plano (TX)

On-site
USD 83,000 - 110,000
Medical, dental, and vision insurance
401(k) with generous match
Tuition Reimbursement
+4
Platform Reliability Engineer — AI Infra & CloudOps
Platform Reliability Engineer — AI Infra & CloudOps

WRITER • New York (NY)

Hybrid
USD 120,000 - 160,000
Generous PTO
Medical, dental, and vision coverage
Paid parental leave (16 weeks)
+5
Software Engineer, Fleet Infra on Hyperscale Azure & Kubernetes
Software Engineer, Fleet Infra on Hyperscale Azure & Kubernetes

Cloudjobs • New York (NY)

On-site
USD 140,000 - 210,000
Relocation assistance
HPC Server Reliability Engineer
HPC Server Reliability Engineer

Advanced Micro Devices, Inc. • Secaucus (NJ)

On-site
USD 140,000 - 190,000
Staff Reliability Engineer, Cloud Infrastructure
Staff Reliability Engineer, Cloud Infrastructure

ZT Systems • Secaucus (NJ)

On-site
USD 105,000 - 154,000
401(k) retirement savings plan
Tuition reimbursement
Medical, dental, and vision coverage
+1
Staff Software Engineer, Compute Control Plane (GPU/CPU)
Staff Software Engineer, Compute Control Plane (GPU/CPU)

Cloudjobs • San Jose (CA)

On-site
USD 180,000 - 280,000
Cash compensation
Equity compensation
Health insurance
+6