AI Infrastructure Network Operations Engineer

OpenAI

San Francisco (CA)

On-site

USD 140,000 - 210,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

OpenAI is seeking an Infrastructure Operations Engineer to operate and improve large-scale Ethernet fabrics supporting GPU clusters, storage, and management infrastructure. The role blends hands-on production operations with automation, observability, and incident response across a global AI network.

You will partner with architecture, deployment, and engineering teams to raise reliability, perform RCA, and automate repetitive tasks with Python to reduce toil and improve efficiency.

Qualifications

  • 5+ years operating large-scale data center/network infrastructure.
  • Experience with high-availability production networks.
  • Hands-on with Layer 2/3, BGP/OSPF/ECMP, and VLANs.
  • Experience with network upgrades, changes, and RCA.

Responsibilities

  • Own the health, availability, and reliability of production AI network infrastructure.
  • Monitor, troubleshoot, and resolve incidents while meeting SLOs.
  • Operate large-scale Ethernet fabrics for GPU compute, storage, and management networks.
  • Execute changes, maintenance, and capacity expansions with minimal impact.
  • Manage hardware lifecycle and preventive maintenance.
  • Support new AI cluster deployments and data center migrations.
  • Partner with CSPs, colocation, and hardware vendors.
  • Perform RCA and drive permanent corrective actions.

Skills

Network operations
Troubleshooting
Automation with Python
Observability & monitoring
On-call readiness

Education

Bachelor’s degree in CS/Engineering
Equivalent practical experience

Tools

Cisco NX-OS
Arista EOS
NVIDIA Spectrum
Juniper JunOS
Prometheus
Grafana
Terraform
Git

Job description

OpenAI is seeking an Infrastructure Operations Engineer to operate and improve large-scale Ethernet fabrics supporting GPU clusters, storage, and management infrastructure. The role blends hands-on production operations with automation, observability, and incident response across a global AI network.

You will partner with architecture, deployment, and engineering teams to raise reliability, perform RCA, and automate repetitive tasks with Python to reduce toil and improve efficiency.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Network Operations Engineer, AI Networking
Network Operations Engineer, AI Networking

OpenAI • San Francisco (CA)

On-site
USD 140,000 - 210,000
AI Infrastructure & Automation Engineer
AI Infrastructure & Automation Engineer

Nscale • New York (NY)

On-site
USD 140,000 - 210,000
Competitive package
Equity
Growth opportunities
Network Engineer: Secure, Scalable & Automated Infra
Network Engineer: Secure, Scalable & Automated Infra

OpenAI • San Francisco (CA)

Hybrid
USD 293,000 - 385,000
Relocation assistance
Hybrid work model (3 days in office)
AI Network Reliability Engineer: Repair & Automation
AI Network Reliability Engineer: Repair & Automation

Meta • Oregon (WI)

On-site
USD 193,000 - 271,000
Bonus
Equity
Benefits
Global HPC Network Engineer for AI Infra
Global HPC Network Engineer for AI Infra

Together • San Francisco (CA)

On-site
USD 190,000 - 280,000
Startup equity
Health insurance
Competitive benefits
AI Infrastructure Repair Network Engineer
AI Infrastructure Repair Network Engineer

Meta • Eagle Mountain (UT)

On-site
USD 193,000 - 271,000
Senior AI Infrastructure Engineer — Scale GPU Clusters Remote
Senior AI Infrastructure Engineer — Scale GPU Clusters Remote

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 280,000 - 420,000
Equity options
Health, vision, dental benefits
Unlimited PTO
+2
AI Infrastructure Engineer - Fleet & Automation
AI Infrastructure Engineer - Fleet & Automation

Nscale • Seattle (WA)

On-site
USD 150,000 - 215,000
Equity
Base salary + equity
Career progression
Senior AI Infrastructure Network Engineer
Senior AI Infrastructure Network Engineer

Nscale • New York (NY)

On-site
USD 150,000 - 240,000
Equity
Comprehensive benefits
Retirement plan
Data Engineer - Infrastructure Analytics at Scale
Data Engineer - Infrastructure Analytics at Scale

Neura Market • San Francisco (CA)

On-site
USD 140,000 - 190,000