Platform Reliability Engineer – Scalable AI Infra

OpenAI

California (MO)

On-site

USD 190,000 - 270,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Relocation assistance

Job summary

OpenAI is seeking experienced reliability engineers to help scale and safeguard our infrastructure at a global level. You will collaborate with researchers, engineers, product managers, and designers to deliver reliable, high‑performing systems for millions of users worldwide.

The role emphasizes building fault‑tolerant designs, implementing SLOs/SLIs, and advancing automation and observability across cloud environments. OpenAI offers relocation assistance for this San Francisco–based position.

Qualifications

  • Bachelor’s degree in Computer Science, Information Technology, or a related field (or equivalent work experience).
  • Proven experience as SWE focused on reliability or a similar role in a fast‑paced, rapidly scaling company.
  • Strong proficiency in cloud infrastructure.
  • Proficiency in programming languages.
  • Experience with containerization technologies and container orchestration platforms like Kubernetes.
  • Knowledge of IaC tools such as Terraform or CloudFormation.
  • Excellent problem‑solving and troubleshooting skills.
  • Strong communication and collaboration skills.
  • Experience with observability tools such as Datadog, Prometheus, Grafana and Splunk.
  • Experience with microservices architecture and service mesh technologies.
  • Knowledge of security best practices in cloud environments.

Responsibilities

  • Design and implement solutions to ensure the scalability of our infrastructure to meet rapidly increasing demands.
  • Build and maintain the load, chaos and synthetic‑testing software leveraged by development teams to make the systems they design and operate more reliable.
  • Build and maintain automation tools to streamline repetitive tasks and improve system reliability.
  • Build and maintain the platform for CPU, storage, GPU, and network lifecycle management to drive efficiency, accountability and dynamic optimization of our resources.
  • Implement fault‑tolerant and resilient design patterns to minimize service disruptions.
  • Develop and maintain service level objectives (SLOs) and service level indicators (SLIs) to measure and ensure system reliability.
  • Partner with researchers, engineers, product managers, and designers to bring new features and research capabilities to the world.
  • Participate in an on‑call rotation to respond to critical incidents and ensure 24/7 system availability.

Skills

Cloud infrastructure
Programming languages
Observability tooling
Microservices
Security best practices
Collaboration
On-call readiness

Education

Bachelor's degree in Computer Science or related field

Tools

Kubernetes
Terraform
CloudFormation

Job description

OpenAI is seeking experienced reliability engineers to help scale and safeguard our infrastructure at a global level. You will collaborate with researchers, engineers, product managers, and designers to deliver reliable, high‑performing systems for millions of users worldwide.

The role emphasizes building fault‑tolerant designs, implementing SLOs/SLIs, and advancing automation and observability across cloud environments. OpenAI offers relocation assistance for this San Francisco–based position.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Infrastructure Reliability Engineer - Distributed Systems
Infrastructure Reliability Engineer - Distributed Systems

OpenAI • San Francisco (CA)

On-site
USD 255,000 - 405,000
Software Engineer – Scalable Distributed Infrastructure
Software Engineer – Scalable Distributed Infrastructure

Neura Market • San Francisco (CA), Northern (KY)

Hybrid
USD 190,000 - 230,000
Resilience Engineer: Build Scalable, Reliable Systems
Resilience Engineer: Build Scalable, Reliable Systems

OpenAI • San Francisco (CA)

On-site
USD 190,000 - 260,000
Relocation assistance
Software Engineer, Resilience Engineering
Software Engineer, Resilience Engineering

OpenAI • California (MO)

On-site
USD 190,000 - 270,000
Relocation assistance
Software Engineer, Resilience Engineering
Software Engineer, Resilience Engineering

OpenAI • San Francisco (CA)

On-site
USD 190,000 - 260,000
Relocation assistance
Data Infrastructure Engineer — Scale & Reliability
Data Infrastructure Engineer — Scale & Reliability

OpenAI • California (MO)

Hybrid
USD 150,000 - 190,000
Relocation assistance
Cloud Infrastructure Engineer - Scale & Security
Cloud Infrastructure Engineer - Scale & Security

OpenAI • California (MO)

On-site
USD 170,000 - 210,000
Senior AI Reliability Engineer - Scale & Resilience
Senior AI Reliability Engineer - Scale & Resilience

Anthropic • San Francisco (CA)

Hybrid
USD 325,000 - 485,000
Competitive compensation
Generous vacation
Flexible working hours
Senior Site Reliability Engineer — AI Platform Scale
Senior Site Reliability Engineer — AI Platform Scale

Future Secure AI • Austin (TX)

On-site
USD 140,000 - 190,000
Senior Security Reliability Engineer – On-Prem & Hybrid Infra
Senior Security Reliability Engineer – On-Prem & Hybrid Infra

OpenAI • California (MO)

On-site
USD 190,000 - 230,000