Resilience Engineer: Scale, Reliability & Automation

OpenAI, Inc.

San Francisco (CA)

On-site

USD 230,000 - 490,000

Full time

11 days ago
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Health Insurance
401(k) Match
Parental Leave
Paid Time Off
Relocation Assistance
Meals in Office

Job summary

OpenAI, Inc. in San Francisco seeks an experienced Software Engineer for Resilience Engineering to ensure reliability, scalability and performance of our rapidly evolving infrastructure.

You will collaborate with researchers, engineers, product managers and designers to bring features to millions while maintaining safety and uptime. The role emphasizes building tooling for load testing, automation, and platform management across CPU, storage, GPU, and network resources, with on-call rotation for

Qualifications

  • Bachelor's degree in Computer Science, Information Technology, or a related field (or equivalent work experience).
  • Proven experience as an SWE focused on reliability or a similar role in a fast-paced, rapidly scaling company.
  • Strong proficiency in cloud infrastructure.
  • Proficiency in programming languages.
  • Experience with containerization technologies and container orchestration platforms like Kubernetes.
  • Knowledge of IaC tools such as Terraform or CloudFormation.
  • Excellent problem-solving and troubleshooting skills.
  • Strong communication and collaboration skills.
  • Experience with observability tools such as Datadog, Prometheus, Grafana and Splunk.
  • Experience with microservices architecture and service mesh technologies.
  • Knowledge of security best practices in cloud environments.

Responsibilities

  • Design and implement solutions to ensure the scalability of our infrastructure to meet rapidly increasing demands.
  • Build and maintain the load, chaos and synthetic-testing software leveraged by development teams to make the systems they design and operate more reliable.
  • Build and maintain automation tools to streamline repetitive tasks and improve system reliability.
  • Build and maintain the platform for CPU, storage, GPU, and network lifecycle management to drive efficiency, accountability and dynamic optimization of our resources.
  • Implement fault-tolerant and resilient design patterns to minimize service disruptions.
  • Develop and maintain service level objectives (SLOs) and service level indicators (SLIs) to measure and ensure system reliability.
  • Partner with researchers, engineers, product managers, and designers to bring new features and research capabilities to the world.
  • Participate in an on-call rotation to respond to critical incidents and ensure 24/7 system availability.

Skills

Cloud infrastructure
Programming languages
Observability
Microservices
Security best practices
IaC
Kubernetes
Cross-functional collaboration

Education

Bachelor's degree in Computer Science or related field

Tools

Kubernetes
Terraform
CloudFormation
Datadog
Prometheus
Grafana
Splunk

Job description

OpenAI, Inc. in San Francisco seeks an experienced Software Engineer for Resilience Engineering to ensure reliability, scalability and performance of our rapidly evolving infrastructure.

You will collaborate with researchers, engineers, product managers and designers to bring features to millions while maintaining safety and uptime. The role emphasizes building tooling for load testing, automation, and platform management across CPU, storage, GPU, and network resources, with on-call rotation for

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Resilience Engineer: Build Scalable, Reliable Systems
Resilience Engineer: Build Scalable, Reliable Systems

OpenAI • San Francisco (CA)

On-site
USD 190,000 - 260,000
Relocation assistance
Software Engineer, Resilience Engineering
Software Engineer, Resilience Engineering

OpenAI • San Francisco (CA)

On-site
USD 190,000 - 260,000
Relocation assistance
Infrastructure Reliability Engineer - Distributed Systems
Infrastructure Reliability Engineer - Distributed Systems

OpenAI • San Francisco (CA)

On-site
USD 255,000 - 405,000
Engineering Manager, Core Platform & Reliability
Engineering Manager, Core Platform & Reliability

OpenAI • San Francisco (CA)

On-site
USD 230,000 - 290,000
Relocation assistance
Senior AI Reliability Engineer - Scale & Resilience
Senior AI Reliability Engineer - Scale & Resilience

Anthropic • San Francisco (CA)

Hybrid
USD 325,000 - 485,000
Competitive compensation
Generous vacation
Flexible working hours
Software Engineer, Resilience Engineering
Software Engineer, Resilience Engineering

OpenAI, Inc. • San Francisco (CA)

On-site
USD 230,000 - 490,000
Health Insurance
401(k) Match
Parental Leave
+3
Site Reliability Engineer — Scale & Resilience for AI Ops
Site Reliability Engineer — Scale & Resilience for AI Ops

HappyRobot • San Francisco (CA)

On-site
USD 120,000 - 160,000
Top-Tier Compensation
Ownership & Autonomy
Opportunity to work at a high-growth startup
Senior SRE: Scale Resilient AI Platforms & Automation
Senior SRE: Scale Resilient AI Platforms & Automation

Relx Plc • Philadelphia

Hybrid
USD 95,000 - 159,000
Senior Resilience Engineer: Scale & Observability
Senior Resilience Engineer: Scale & Observability

Persona • San Francisco (CA)

On-site
USD 130,000 - 180,000
Medical benefits
Dental benefits
Vision benefits
+6
Software Engineer, GenAI Resilience for AWS Infra
Software Engineer, GenAI Resilience for AWS Infra

Amazon Web Services (AWS) • Portland (OR)

On-site
USD 144,000 - 194,000