Software Engineer, Compute Infrastructure

CV in

Northern (KY)

Hybrid

USD 180,000 - 240,000

Full time

7 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

OpenAI is seeking a Software Engineer for Compute Infrastructure to build and optimize the platform that powers frontier AI. You will design, provision, and operate large-scale systems connecting GPUs, CPUs, networks, storage, and orchestration to run demanding workloads.

Ideal candidates have distributed systems experience, HPC familiarity, and a track record of reliability engineering, observability, and performance tuning at scale.

Qualifications

  • Strong software engineering skills and experience building, operating, or improving production infrastructure systems are essential.
  • Experience in distributed systems, operating systems, networking protocols, RDMA, NCCL or collective communication, storage, Kubernetes, scheduling, observability, reliability engineering, high‑performance computing, GPU infrastructure, CaaS, agent infrastructure, hardware‑aware performance optimization, benchmarking, developer experience, or infrastructure tooling is required.

Responsibilities

  • You will build and deeply optimize reliable system software for large-scale compute systems that run AI workloads.
  • You will design and operate infrastructure across accelerators, CPUs, NICs, switches, networking protocols, storage, data centers, cluster orchestration, scheduling, and fleet health.
  • Profiling, benchmarking, and optimization of training workloads across compute, memory, storage, networking, NCCL and collective communication will form part of your daily work.

Skills

Distributed systems
High-performance computing
Observability
Reliability engineering

Tools

Kubernetes
NCCL/GPUs
RDMA
Observability tooling

Job description

Role Overview

The role belongs to a Software Engineer focused on Compute Infrastructure. Candidates joining this position build the platform that converts massive compute capacity into a dependable engine for frontier AI. This position designs, provisions, schedules, operates, and optimizes the systems that connect accelerators, CPUs, networks, storage, data centers, orchestration software, agent infrastructure, developer tools, and observability into one coherent experience for researchers and product teams.

Scope of Impact

This role operates across the entire Compute Infrastructure stack. Responsibilities include capacity planning and cluster lifecycle, bare-metal automation, distributed systems, Kubernetes and scheduling, deep system optimization, high-performance networking, storage, fleet health, reliability, workload profiling, benchmarking, and the developer experience. At this scale, improvements in communication, scheduling, hardware efficiency, or debugging workflows can compound into meaningful research velocity. The company hires across Compute Infrastructure rather than for a single narrow team, using this opening to match strong engineers to problems where they can generate the most leverage.

Potential Work Areas

Candidates may work close to hardware or close to users, on CaaS and agent infrastructure, or on the control planes and data planes in between. You might help bring new supercomputing capacity online, optimize training workloads from profiler traces and benchmarks, improve NCCL and collective communication behavior, reason about GPUs, NICs, topology, firmware, thermals, and failure modes, or design abstractions that make heterogeneous clusters feel like one coherent platform.

Expected Contributions

You will build and deeply optimize reliable system software for large-scale compute systems that run some of the world's most demanding AI workloads. You will design and operate infrastructure across accelerators, CPUs, NICs, switches, networking protocols, storage, data centers, cluster orchestration, scheduling, and fleet health. Profiling, benchmarking, and optimization of training workloads across compute, memory, storage, networking, NCCL and collective communication, and cluster scheduling bottlenecks will form part of your daily work. You will create hardware‑aware automation that makes provisioning, firmware and driver upgrades, incident response, and day‑to‑day operations faster and less error‑prone. Building CaaS, agent infrastructure, profiling, observability, benchmarking, and platform tools that help researchers, product engineers, and operators launch, debug, and optimize workloads with less friction will be central. You will turn operational lessons into better systems, stronger abstractions, and clearer ownership boundaries across teams. Collaboration across research, engineering, security, networking, hardware, and data center teams will help make compute capacity more capable and easier to use.

Ideal Candidate Profile

You have built or operated distributed systems, infrastructure platforms, high‑performance computing environments, large‑scale networking systems, Kubernetes clusters, developer tools, or production systems with demanding reliability requirements. You enjoy working across layers of the stack and are comfortable moving between software, hardware, networking, systems performance, reliability, and user needs. You care about making complex infrastructure understandable, observable, and usable for the people depending on it. You can diagnose hard problems under real operational pressure while still investing in long‑term engineering quality. You like building leverage for others, whether through APIs, automation, debugging tools, CaaS and agent infrastructure primitives, workflow improvements, or better platform abstractions. You are motivated by scale, efficiency, reliability, and disciplined measurement through benchmarks, profiles, and production evidence. You communicate clearly, take ownership, and work well with teams whose constraints and goals differ from your own.

Required Qualifications

Strong software engineering skills and experience building, operating, or improving production infrastructure systems are essential. Experience in one or more relevant areas such as distributed systems, operating systems, networking protocols, RDMA, NCCL or collective communication, storage, Kubernetes, scheduling, observability, reliability engineering, high‑performance computing, GPU infrastructure, CaaS, agent infrastructure, hardware‑aware performance optimization, benchmarking, developer experience, or infrastructure tooling is required. You must be able to debug complex system behavior across software, hardware, networking, and workload layers, then turn findings into robust improvements. Comfort with ambiguity, strong ownership, and a bias toward practical, durable solutions are necessary. A genuine interest in working on infrastructure that directly enables frontier AI research and product impact is required.

About OpenAI

OpenAI is an AI research and deployment company dedicated to ensuring that general‑purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve our mission, we must encompass and value the many different perspectives, voices, and experiences that form the full spectrum of humanity. OpenAI is an equal opportunity employer, and we do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, age, veteran status, disability, genetic information, or other applicable legally protected characteristic. For additional information, please see OpenAI's https://cdn.OpenAI. com/policies/eeo-policy-statement.pdf. Background checks for applicants will be administered in accordance with applicable law, and qualified applicants with arrest or conviction records will be considered for employment consistent with those laws, including the San Francisco Fair Chance Ordinance, the Los Angeles County Fair Chance Ordinance for Employers, and the California Fair Chance Act, for US-based candidates. For unincorporated Los Angeles County workers: we reasonably believe that criminal history may have a direct, adverse and negative relationship with the following job duties, potentially resulting in the withdrawal of a conditional offer of employment: protect computer hardware entrusted to you from theft, loss or damage; return all computer hardware in your possession (including the data contained therein) upon termination of employment or end of assignment; and maintain the confidentiality of proprietary, confidential, and non‑public information. In addition, job duties require access to secure and protected information technology systems and related data security obligations. To notify OpenAI that you believe this job posting is non‑compliant, please submit a report through this form https://form.asana.com/?d=57088692298241&k=5MqR40fZd7jlxVUh5J-UeA. No response will be provided to inquiries unrelated to job posting compliance. We are committed to providing reasonable accommodations to applicants with disabilities, and requests can be made via this link https://form.asana.com/ ? k=bQ7w9h3iexRlicUdWRiwvg&d=57088692298241. OpenAI Global Applicant Privacy Policy https:// cdn.openai.com/policies/global-employee-and-contractor-privacy-policy.pdf At OpenAI, we believe artificial intelligence has the potential to help people solve immense global challenges, and we want the upside of AI to be widely shared. Join us in shaping the future of technology.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineer - Data Aquisition (systems)
Software Engineer - Data Aquisition (systems)

Triwill Group • San Francisco (CA)

On-site
USD 180,000 - 260,000
Software Engineer, Compute Infrastructure
Software Engineer, Compute Infrastructure

OpenAI • Los Angeles (CA)

On-site
USD 230,000 - 405,000
Equity
Flexible work environment
Health benefits
CPU/Storage/PoP-WAN Program Manager OpenAI San Francisco
CPU/Storage/PoP-WAN Program Manager OpenAI San Francisco

Neura Market • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Software Engineer, GPT Infrastructure
Software Engineer, GPT Infrastructure

OpenAI • Seattle (WA)

On-site
USD 180,000 - 240,000
Software Engineer, GPT Infrastructure
Software Engineer, GPT Infrastructure

Slope • San Francisco (CA)

On-site
USD 180,000 - 260,000
Software Engineer - Data Aquisition (systems)
Software Engineer - Data Aquisition (systems)

OpenAI • San Francisco (CA)

On-site
USD 255,000 - 405,000
Equity
Software Engineer, Workload Enablement
Software Engineer, Workload Enablement

OpenAI • Seattle (WA)

On-site
USD 293,000 - 455,000
Data Engineer, Scaling Analytics
Data Engineer, Scaling Analytics

OpenAI • San Francisco (CA)

On-site
USD 120,000 - 150,000
Data Engineer, CPU & Storage
Data Engineer, CPU & Storage

OpenAI • San Francisco (CA)

On-site
USD 140,000 - 210,000
Software Engineer, Platform Systems
Software Engineer, Platform Systems

OpenAI • San Francisco (CA)

On-site
USD 310,000 - 460,000