Resilience Engineer: Build Scalable, Reliable Systems

OpenAI

San Francisco (CA)

On-site

USD 190,000 - 260,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Relocation assistance

Job summary

OpenAI is hiring a reliability-focused SWE to scale and protect our infrastructure as the company grows. You will build and maintain tools for load testing, chaos experiments, and automation to reduce incidents and improve performance across services.

You will collaborate with research, product, and design teams to deploy reliable features and manage complex distributed systems with strong security and observability practices.

Qualifications

  • Experience as SWE focused on reliability in a fast-paced, scaling company.
  • Strong knowledge of cloud infrastructure and distributed systems.
  • Proficiency with containerization and orchestration (Kubernetes).
  • Experience with observability tooling and incident response.
  • Familiarity with IaC and security best practices.

Responsibilities

  • Design scalable infrastructure solutions to meet rapidly increasing demand.
  • Develop and maintain load, chaos, and synthetic-testing tooling to improve reliability.
  • Create automation to streamline repetitive tasks and improve system reliability.
  • Manage lifecycle of CPU, storage, GPU, and network resources for efficiency.
  • Implement fault-tolerant patterns and define SLOs/SLIs for reliability metrics.
  • Collaborate with researchers, engineers, and product teams to deploy reliable features.
  • Participate in on-call rotation to ensure 24/7 system availability.

Skills

Reliability engineering
Cloud infrastructure
Observability
Infrastructure as Code
Kubernetes
Security practices
Cross-functional collaboration
SRE tooling

Education

Bachelor's degree

Tools

Kubernetes
Terraform
CloudFormation
Datadog
Prometheus
Grafana
Splunk
Docker

Job description

OpenAI is hiring a reliability-focused SWE to scale and protect our infrastructure as the company grows. You will build and maintain tools for load testing, chaos experiments, and automation to reduce incidents and improve performance across services.

You will collaborate with research, product, and design teams to deploy reliable features and manage complex distributed systems with strong security and observability practices.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineer, Resilience Engineering
Software Engineer, Resilience Engineering

OpenAI • San Francisco (CA)

On-site
USD 190,000 - 260,000
Relocation assistance
Infrastructure Reliability Engineer - Distributed Systems
Infrastructure Reliability Engineer - Distributed Systems

OpenAI • San Francisco (CA)

On-site
USD 255,000 - 405,000
Resilience Engineering Manager — Production Reliability
Resilience Engineering Manager — Production Reliability

Affirm • Salt Lake City (UT)

On-site
USD 210,000 - 260,000
Health coverage
Equity rewards
Tech stipends
+1
Senior AI Reliability Engineer - Scale & Resilience
Senior AI Reliability Engineer - Scale & Resilience

Anthropic • San Francisco (CA)

Hybrid
USD 325,000 - 485,000
Competitive compensation
Generous vacation
Flexible working hours
Engineering Manager, Resilience & Chaos Engineering
Engineering Manager, Resilience & Chaos Engineering

Affirm • San Diego (CA)

On-site
USD 230,000 - 290,000
Health care coverage
Flexible Spending Wallets
Time off
+1
Remote Engineering Manager, Resilience & Chaos Engineering
Remote Engineering Manager, Resilience & Chaos Engineering

Affirm • Detroit (MI)

On-site
USD 204,000 - 290,000
Health coverage
Equity
Remote-friendly culture
Senior Site Reliability Engineer — AI Platform Scale
Senior Site Reliability Engineer — AI Platform Scale

Future Secure AI • Austin (TX)

On-site
USD 140,000 - 190,000
Site Reliability Engineer — Scale & Resilience for AI Ops
Site Reliability Engineer — Scale & Resilience for AI Ops

HappyRobot • San Francisco (CA)

On-site
USD 120,000 - 160,000
Senior Infrastructure Engineer for OpenAI Platform (Equity)
Senior Infrastructure Engineer for OpenAI Platform (Equity)

OpenAI • Bellevue (WA)

Hybrid
USD 293,000 - 325,000
Relocation assistance
Senior Site Reliability Engineer - Drive Resilient Systems
Senior Site Reliability Engineer - Drive Resilient Systems

United States Digital Space LLC • United States

Remote
USD 83,000 - 115,000
Health care coverage
Flexible Spending Wallets
Time off
+3