Infrastructure Reliability Engineer - Distributed Systems

OpenAI

San Francisco (CA)

On-site

USD 255,000 - 405,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

OpenAI is hiring software engineers for their Infrastructure team in San Francisco. This role focuses on scaling and securing critical systems that power AI applications such as ChatGPT. Candidates will tackle complex technical challenges and collaborate with various teams to enhance system resilience.

The ideal candidate has over 4 years of experience in distributed systems, proficiency with cloud platforms, and a passion for engineering excellence.

Qualifications

  • 4+ years of relevant industry experience.
  • Proven experience as a reliability engineer or similar role.
  • Strong proficiency in cloud infrastructure and IaC tools.

Responsibilities

  • Design, build, and operate reliable systems.
  • Identify and fix performance bottlenecks.
  • Contribute to incident response and best practices.

Skills

Distributed systems principles
Performance optimization
Container orchestration (Kubernetes)
Linux environments
Cloud platforms (AWS, GCP, Azure)
Observability tools (Datadog, Prometheus, etc.)
Microservices architecture
Security best practices

Education

4+ years of relevant industry experience
Experience leading complex projects

Tools

Terraform
CI/CD pipelines

Job description

OpenAI is hiring software engineers for their Infrastructure team in San Francisco. This role focuses on scaling and securing critical systems that power AI applications such as ChatGPT. Candidates will tackle complex technical challenges and collaborate with various teams to enhance system resilience.

The ideal candidate has over 4 years of experience in distributed systems, proficiency with cloud platforms, and a passion for engineering excellence.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Infrastructure Engineer — Distributed Systems
Senior Infrastructure Engineer — Distributed Systems

OpenAI • San Francisco (CA)

On-site
USD 180,000 - 280,000
Resilience Engineer: Scale, Reliability & Automation
Resilience Engineer: Scale, Reliability & Automation

OpenAI, Inc. • San Francisco (CA)

On-site
USD 230,000 - 490,000
Health Insurance
401(k) Match
Parental Leave
+3
Resilience Engineer: Build Scalable, Reliable Systems
Resilience Engineer: Build Scalable, Reliable Systems

OpenAI • San Francisco (CA)

On-site
USD 190,000 - 260,000
Relocation assistance
Backend Engineer - Distributed Systems & Secure APIs
Backend Engineer - Distributed Systems & Secure APIs

OpenAI • San Francisco (CA)

On-site
USD 100,000 - 150,000
Engineering Manager, Core Platform & Reliability
Engineering Manager, Core Platform & Reliability

OpenAI • San Francisco (CA)

On-site
USD 230,000 - 290,000
Relocation assistance
Backend Engineer - Distributed Systems (Scale & Reliability)
Backend Engineer - Distributed Systems (Scale & Reliability)

EngineersOfAI • San Francisco (CA)

On-site
USD 120,000 - 150,000
Software Engineer, Resilience Engineering
Software Engineer, Resilience Engineering

OpenAI • San Francisco (CA)

On-site
USD 190,000 - 260,000
Relocation assistance
Software Engineer, Infrastructure
Software Engineer, Infrastructure

OpenAI • San Francisco (CA)

On-site
USD 180,000 - 280,000
Platform Reliability Engineer — AI Infra & CloudOps
Platform Reliability Engineer — AI Infra & CloudOps

WRITER • New York (NY)

Hybrid
USD 120,000 - 160,000
Generous PTO
Medical, dental, and vision coverage
Paid parental leave (16 weeks)
+5
Software Engineer, Cloud Infrastructure
Software Engineer, Cloud Infrastructure

OpenAI, Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 255,000 - 490,000