Software Engineer, DevOps

Engg

San Francisco (CA)

On-site

USD 180,000 - 270,000

Full time

6 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

OpenAI’s Hardware organization seeks an engineer to define and build scalable CI/CD and GitOps infrastructure for frontier models running on OpenAI’s custom accelerators. You will work across model workloads, accelerator software, systems infrastructure, and developer productivity, shaping monorepo architecture, test strategies, and performance baselines to accelerate safe production readiness.

You will collaborate with model, compiler, kernel, runtime, firmware, validation, and hardware teams

Qualifications

  • Experience building and operating CI/CD or production engineering systems.
  • Familiarity with GitOps, release promotion, rollback, reproducibility, and automation policies.
  • Ability to design monorepos, build systems, dependency graphs, and change-impact analysis.

Responsibilities

  • Define the model CI strategy for frontier models running on OpenAI custom silicon and integrated accelerator systems.
  • Build production CI/CD pipelines for functional regression, performance benchmarking, software stability, and release qualification.
  • Design representative test matrices across models, configurations, hardware generations, software components, and deployment environments.
  • Create reliable performance baselines, regression detection, bisect and triage workflows, and clear ownership for failures.
  • Establish GitOps practices for reproducible configuration, promotion, rollback, auditability, and environment consistency.
  • Shape monorepo architecture, dependency management, build and test boundaries, change validation, and developer workflows at scale.
  • Develop scalable orchestration, artifact management, caching, scheduling, observability, and capacity controls for accelerator-backed CI.
  • Partner with model, compiler, kernel, runtime, firmware, validation, and hardware teams to translate release risks into automated gates.
  • Improve CI reliability, speed, debuggability, and cost efficiency while maintaining high confidence in production software.

Skills

CI/CD pipelines
GitOps
Python
Go
Rust
Distributed systems

Tools

Git
Kubernetes
CI tools

Job description

About the Team

OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform.

About the Role

You will define and build the CI/CD systems that qualify frontier models and the software stack running on OpenAI’s custom AI accelerators. Your work will make correctness, performance, and stability measurable on every change, from individual commits through production releases. This role spans model workloads, accelerator software, systems infrastructure, and developer productivity. You will establish scalable CI/CD, GitOps, and monorepo practices; design trustworthy regression and benchmarking pipelines; and partner with model, compiler, kernel, runtime, firmware, and hardware teams to shorten feedback loops without compromising production readiness.

In this role, you will:

  • Define the model CI strategy for frontier models running on OpenAI custom silicon and integrated accelerator systems.
  • Build production CI/CD pipelines for functional regression, performance benchmarking, software stability, and release qualification.
  • Design representative test matrices across models, configurations, hardware generations, software components, and deployment environments.
  • Create reliable performance baselines, regression detection, bisect and triage workflows, and clear ownership for failures.
  • Establish GitOps practices for reproducible configuration, promotion, rollback, auditability, and environment consistency.
  • Shape monorepo architecture, dependency management, build and test boundaries, change validation, and developer workflows at scale.
  • Develop scalable orchestration, artifact management, caching, scheduling, observability, and capacity controls for accelerator-backed CI.
  • Partner with model, compiler, kernel, runtime, firmware, validation, and hardware teams to translate release risks into automated gates.
  • Improve CI reliability, speed, debuggability, and cost efficiency while maintaining high confidence in production software.

You might thrive in this role if:

  • Have built or operated large-scale CI/CD, developer infrastructure, test automation, or production engineering systems.
  • Understand modern software delivery practices, including GitOps, release promotion, rollback, reproducibility, and policy-driven automation.
  • Have experience designing monorepos, build systems, dependency graphs, test selection, or change-impact analysis.
  • Can design regression and performance-benchmarking systems with stable baselines, useful metrics, and actionable failure diagnosis.
  • Are proficient in Python, Go, Rust, C++, or another language used to build reliable infrastructure and automation.
  • Have worked with distributed systems, schedulers, containers, clusters, or heterogeneous compute infrastructure.
  • Can collaborate with model and systems engineers to turn complex accelerator workloads into repeatable production qualification.
  • Care deeply about reliability, observability, developer experience, and reducing time from code change to trustworthy signal.

To comply with U.S. export control laws and regulations, candidates for this role may need to meet certain legal status requirements as provided in those laws and regulations.

About OpenAI

OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve our mission, we must encompass and value the many different perspectives, voices, and experiences that form the full spectrum of humanity.

We are an equal opportunity employer, and we do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, age, veteran status, disability, genetic information, or other applicable legally protected characteristic.

For additional information, please see OpenAI’s Affinity statement https://cdn.openai.com/policies/eeo-policy-statement.pdf.

Background checks for applicants will be administered in accordance with applicable law, and qualified applicants with arrest or conviction records will be considered for employment consistent with those laws, including the San Francisco Fair Chance Ordinance, the Los Angeles County Fair Chance Ordinance for Employers, and the California Fair Chance Act, for US-based candidates.

For unincorporated Los Angeles County workers: we reasonably believe that criminal history may have a direct, adverse and negative relationship with the following job duties, potentially resulting in the withdrawal of a conditional offer of employment: protect computer hardware entrusted to you from theft, loss or damage; return all computer hardware in your possession (including the data contained therein) upon termination of employment or end of assignment; and maintain the confidentiality of proprietary, confidential, and non-public information.

In addition, job duties require access to secure and protected information technology systems and related data security obligations.

To notify OpenAI that you believe this job posting is non-compliant, please submit a report through this form https://form.asana.com/?d=57018692298241&k=5MqR40fZd7jlxVUh5J-UeA. No response will be provided to inquiries unrelated to job posting compliance.

We are committed to providing reasonable accommodations to applicants with disabilities, and requests can be made vi

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Software Engineer, DevOps
Software Engineer, DevOps

Xapply • San Francisco (CA)

On-site
USD 180,000 - 240,000
Software Engineer, DevOps
Software Engineer, DevOps

OpenAI • San Francisco (CA)

On-site
USD 177,000 - 327,000
Systems Software Engineer, Silicon Bringup
Systems Software Engineer, Silicon Bringup

Triwill Group • United States

On-site
USD 180,000 - 230,000
Software Engineer, Model Runtime
Software Engineer, Model Runtime

Triwill Group • United States

On-site
USD 180,000 - 240,000
Software Engineer, GPT Infrastructure
Software Engineer, GPT Infrastructure

Slope • San Francisco (CA)

On-site
USD 180,000 - 260,000
Software Engineer, GPT Infrastructure
Software Engineer, GPT Infrastructure

OpenAI • Seattle (WA)

On-site
USD 180,000 - 240,000
AI Accelerator Systems Engineer — Scale & Performance
AI Accelerator Systems Engineer — Scale & Performance

OpenAI • San Francisco (CA)

On-site
USD 120,000 - 180,000
Software Engineer, AI accelerator Runtime
Software Engineer, AI accelerator Runtime

Triwill Group • United States

On-site
USD 180,000 - 240,000
Software Engineer, DevOps
Software Engineer, DevOps

OpenAI, Inc. • San Francisco (CA)

Hybrid
USD 177,000 - 327,000
Medical, dental, and vision insurance
401(k) retirement plan with employer 1
Paid parental leave
+4
Software Engineer, Build Systems / CI
Software Engineer, Build Systems / CI

jobr.pro • San Francisco (CA)

On-site
USD 150,000 - 230,000