Software Engineer, DevOps

OpenAI

San Francisco (CA)

On-site

USD 177,000 - 327,000

Full time

22 hours ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

OpenAI is seeking a seasoned CI/CD/DevOps engineer to define and build the CI/CD systems that qualify frontier models and the software stack running on our custom AI accelerators. Your work will make correctness, performance, and stability measurable on every change, from commits to production releases.

This role spans model workloads, accelerator software, systems infrastructure, and developer productivity.

Qualifications

  • Built or operated large-scale CI/CD, developer infrastructure, test automation, or production engineering systems.
  • Understand modern software delivery practices, including GitOps, release promotion, rollback, reproducibility, and policy-driven automation.
  • Have experience designing monorepos, build systems, dependency graphs, test selection, or change-impact analysis.
  • Can design regression and performance-benchmarking systems with stable baselines, useful metrics, and actionable failure diagnosis.
  • Are proficient in Python, Go, Rust, C++, or another language used to build reliable infrastructure and automation.

Responsibilities

  • Define the model CI strategy for frontier models running on OpenAI custom silicon and integrated accelerator systems.
  • Build production CI/CD pipelines for functional regression, performance benchmarking, software stability, and release qualification.
  • Design representative test matrices across models, configurations, hardware generations, software components, and deployment environments.
  • Create reliable performance baselines, regression detection, bisect and triage workflows, and clear ownership for failures.
  • Establish GitOps practices for reproducible configuration, promotion, rollback, auditability, and environment consistency.
  • Shape monorepo architecture, dependency management, build and test boundaries, change validation, and developer workflows at scale.
  • Develop scalable orchestration, artifact management, caching, scheduling, observability, and capacity controls for accelerator-backed CI.
  • Partner with model, compiler, kernel, runtime, firmware, validation, and hardware teams to translate release risks into automated gates.
  • Improve CI reliability, speed, debuggability, and cost efficiency while maintaining high confidence in production software.

Skills

CI/CD
Python
Go
Rust
C++
Distributed systems

Tools

GitOps
Monorepo
Kubernetes

Job description

About The Team

OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform.

About The Team

OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform.

About The Role

You will define and build the CI/CD systems that qualify frontier models and the software stack running on OpenAI’s custom AI accelerators. Your work will make correctness, performance, and stability measurable on every change, from individual commits through production releases. This role spans model workloads, accelerator software, systems infrastructure, and developer productivity. You will establish scalable CI/CD, GitOps, and monorepo practices; design trustworthy regression and benchmarking pipelines; and partner with model, compiler, kernel, runtime, firmware, and hardware teams to shorten feedback loops without compromising production readiness.

In This Role, You Will
  • Define the model CI strategy for frontier models running on OpenAI custom silicon and integrated accelerator systems.
  • Build production CI/CD pipelines for functional regression, performance benchmarking, software stability, and release qualification.
  • Design representative test matrices across models, configurations, hardware generations, software components, and deployment environments.
  • Create reliable performance baselines, regression detection, bisect and triage workflows, and clear ownership for failures.
  • Establish GitOps practices for reproducible configuration, promotion, rollback, auditability, and environment consistency.
  • Shape monorepo architecture, dependency management, build and test boundaries, change validation, and developer workflows at scale.
  • Develop scalable orchestration, artifact management, caching, scheduling, observability, and capacity controls for accelerator-backed CI.
  • Partner with model, compiler, kernel, runtime, firmware, validation, and hardware teams to translate release risks into automated gates.
  • Improve CI reliability, speed, debuggability, and cost efficiency while maintaining high confidence in production software.
You Might Thrive In This Role If
  • Have built or operated large-scale CI/CD, developer infrastructure, test automation, or production engineering systems.
  • Understand modern software delivery practices, including GitOps, release promotion, rollback, reproducibility, and policy-driven automation.
  • Have experience designing monorepos, build systems, dependency graphs, test selection, or change-impact analysis.
  • Can design regression and performance-benchmarking systems with stable baselines, useful metrics, and actionable failure diagnosis.
  • Are proficient in Python, Go, Rust, C++, or another language used to build reliable infrastructure and automation.
  • Have worked with distributed systems, schedulers, containers, clusters, or heterogeneous compute infrastructure.
  • Can collaborate with model and systems engineers to turn complex accelerator workloads into repeatable production qualification.
  • Care deeply about reliability, observability, developer experience, and reducing time from code change to trustworthy signal.

To comply with U.S. export control laws and regulations, candidates for this role may need to meet certain legal status requirements as provided in those laws and regulations.

About OpenAI

OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve our mission, we must encompass and value the many different perspectives, voices, and experiences that form the full spectrum of humanity. We are an equal opportunity employer, and we do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, age, veteran status, disability, genetic information, or other applicable legally protected characteristic. For additional information, please see OpenAI’s Affiantive Action and Equal Employment Opportunity Policy Statement. Background checks for applicants will be administered in accordance with applicable law, and qualified applicants with arrest or conviction records will be considered for employment consistent with those laws, including the San Francisco Fair Chance Ordinance, the Los Angeles County Fair Chance Ordinance for Employers, and the California Fair Chance Act, for US-based candidates. For unincorporated Los Angeles County workers: we reasonably believe that criminal history may have a direct, adverse and negative relationship with the following job duties, potentially resulting in the withdrawal of a conditional offer of employment: protect computer hardware entrusted to you from theft, loss or damage; return all computer hardware in your possession (including the data contained therein) upon termination of employment or end of assignment; and maintain the confidentiality of proprietary, confidential, and non-public information. In addition, job duties require access to secure and protected information technology systems and related data security obligations. To notify OpenAI that you believe this job posting is non-compliant, please submit a report through this form. No response will be provided to inquiries unrelated to job posting compliance. We are committed to providing reasonable accommodations to applicants with disabilities, and requests can be made via this link. OpenAI Global Applicant Privacy Policy At OpenAI, we believe artificial intelligence has the potential to help people solve immense global challenges, and we want the upside of AI to be widely shared. Join us in shaping the future of technology.

Compensation Range: $177K - $327K

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Software Engineer, GPT Infrastructure
Software Engineer, GPT Infrastructure

OpenAI • San Francisco (CA)

On-site
USD 293,000 - 385,000
Software Engineer, DevOps
Software Engineer, DevOps

Xapply • San Francisco (CA)

On-site
USD 180,000 - 240,000
Software Engineer, Model Runtime
Software Engineer, Model Runtime

OpenAI, Inc. • San Francisco (CA)

On-site
USD 266,000 - 445,000
Equity
Health insurance
401(k) match
+2
Software Engineer, AI accelerator Runtime
Software Engineer, AI accelerator Runtime

OpenAI • San Francisco (CA)

On-site
USD 266,000 - 445,000
Software Engineer, Hardware
Software Engineer, Hardware

Slope • San Francisco (CA)

On-site
USD 310,000 - 460,000
Technical Program Manager, AI Accelerator Software
Technical Program Manager, AI Accelerator Software

OpenAI • San Francisco (CA)

Hybrid
USD 302,000 - 445,000
Relocation assistance
Hybrid work model
Software Engineer, Model Runtime
Software Engineer, Model Runtime

OpenAI • San Francisco (CA)

On-site
USD 266,000 - 445,000
AI Accelerator Systems Engineer — Scale & Performance
AI Accelerator Systems Engineer — Scale & Performance

OpenAI • San Francisco (CA)

On-site
USD 120,000 - 180,000
Systems Software Engineer, Silicon Bringup
Systems Software Engineer, Silicon Bringup

OpenAI • San Francisco (CA)

On-site
USD 266,000 - 445,000
Software Engineer, Model Runtime
Software Engineer, Model Runtime

Triwill Group • United States

On-site
USD 180,000 - 240,000