Infrastructure Reliability Engineer - Distributed Systems

OpenAI

San Francisco (CA)

On-site

USD 255,000 - 405,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

OpenAI is hiring software engineers for their Infrastructure team in San Francisco. This role focuses on scaling and securing critical systems that power AI applications such as ChatGPT. Candidates will tackle complex technical challenges and collaborate with various teams to enhance system resilience.

The ideal candidate has over 4 years of experience in distributed systems, proficiency with cloud platforms, and a passion for engineering excellence.

Qualifications

  • 4+ years of relevant industry experience.
  • Proven experience as a reliability engineer or similar role.
  • Strong proficiency in cloud infrastructure and IaC tools.

Responsibilities

  • Design, build, and operate reliable systems.
  • Identify and fix performance bottlenecks.
  • Contribute to incident response and best practices.

Skills

Distributed systems principles
Performance optimization
Container orchestration (Kubernetes)
Linux environments
Cloud platforms (AWS, GCP, Azure)
Observability tools (Datadog, Prometheus, etc.)
Microservices architecture
Security best practices

Education

4+ years of relevant industry experience
Experience leading complex projects

Tools

Terraform
CI/CD pipelines

Job description

OpenAI is hiring software engineers for their Infrastructure team in San Francisco. This role focuses on scaling and securing critical systems that power AI applications such as ChatGPT. Candidates will tackle complex technical challenges and collaborate with various teams to enhance system resilience.

The ideal candidate has over 4 years of experience in distributed systems, proficiency with cloud platforms, and a passion for engineering excellence.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Backend Engineer - Distributed Systems & Secure APIs
Backend Engineer - Distributed Systems & Secure APIs

OpenAI • San Francisco (CA)

On-site
USD 100,000 - 150,000
Resilience Engineer: Build Scalable, Reliable Systems
Resilience Engineer: Build Scalable, Reliable Systems

OpenAI • San Francisco (CA)

On-site
USD 190,000 - 260,000
Relocation assistance
Robotics Data Infrastructure Engineer (Distributed Systems)
Robotics Data Infrastructure Engineer (Distributed Systems)

OpenAI • Los Angeles (CA)

Hybrid
USD 230,000 - 385,000
Senior Infrastructure Engineer - Rust/C++ (Distributed Systems)
Senior Infrastructure Engineer - Rust/C++ (Distributed Systems)

Neura Market • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Relocation assistance
Hybrid work model
Data Acquisition Systems Engineer — Infra & Reliability
Data Acquisition Systems Engineer — Infra & Reliability

OpenAI • San Francisco (CA)

On-site
USD 255,000 - 405,000
Equity
Backend Engineer - Distributed Systems (Scale & Reliability)
Backend Engineer - Distributed Systems (Scale & Reliability)

EngineersOfAI • San Francisco (CA)

On-site
USD 120,000 - 150,000
Software Engineer, Resilience Engineering
Software Engineer, Resilience Engineering

OpenAI • San Francisco (CA)

On-site
USD 190,000 - 260,000
Relocation assistance
Senior AI Infrastructure Engineer — Scale, Reliability & Automation
Senior AI Infrastructure Engineer — Scale, Reliability & Automation

AI Chopping Block • San Francisco (CA)

On-site
USD 190,000 - 270,000
Equity
Health insurance
Startup benefits
Machine Learning Engineer, Distributed Data Systems - Robotics OpenAI San Francisco
Machine Learning Engineer, Distributed Data Systems - Robotics OpenAI San Francisco

Neura Market • San Francisco (CA), Northern (KY)

Hybrid
USD 140,000 - 210,000
Relocation assistance
Platform Reliability Engineer — AI Infra & CloudOps
Platform Reliability Engineer — AI Infra & CloudOps

WRITER • New York (NY)

Hybrid
USD 120,000 - 160,000
Generous PTO
Medical, dental, and vision coverage
Paid parental leave (16 weeks)
+5