AI Network Operations Engineer (Production & Automation)

Precision Labs

San Francisco, Northern (CA, KY)

Hybrid

USD 150,000 - 210,000

Full time

13 days ago
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

OpenAI is seeking an Infrastructure Operations Engineer to run and improve the large-scale Ethernet fabrics that support GPU clusters, storage, and management networks. You will blend hands-on production operations with automation, observability, and incident response across a global AI network.

The ideal candidate has 5+ years in high-availability data center or cloud network operations and can handle physical-layer troubleshooting as well as routing, fabric behavior, and RCA.

Qualifications

  • 5+ years operating large-scale data center, cloud, AI, or HPC network infrastructure.
  • Experience supporting production network environments with high-availability requirements.
  • Hands-on experience with Cisco NX‑OS, Arista EOS, NVIDIA Spectrum / Cumulus Linux, or Juniper JunOS.
  • Strong knowledge of Layer 2 and Layer 3 networking, BGP, OSPF, ECMP, MLAG, LACP, VRFs, and VLANs.
  • Excellent analytical and troubleshooting skills, with cross-team communication.

Responsibilities

  • Own the operational health, availability, and reliability of production AI network infrastructure across data centers.
  • Monitor, troubleshoot, and resolve network incidents; meet SLOs and minimize MTTR.
  • Operate and maintain large-scale Ethernet fabrics supporting compute, storage, and management networks.
  • Execute production network changes, maintenance windows, and capacity expansions with minimal impact.
  • Manage hardware lifecycle, including switch and optics replacements, RMA coordination, and upgrades.
  • Support new AI cluster deployments, data center expansions, and migrations with deployment teams.
  • Partner with CSPs, colocation providers, and hardware vendors to maintain production infrastructure.
  • Perform RCA for incidents and drive permanent corrective actions.
  • Build and maintain monitoring, telemetry, dashboards, and alerting for observability.
  • Develop and improve runbooks, playbooks, and standard operating procedures.
  • Automate repetitive tasks using Python and infrastructure automation frameworks.

Skills

Networking expertise
Incident response
Automation with Python
SRE mindset
Cross-team communication

Education

Bachelor's degree in Computer Science or related field

Tools

Cisco NX-OS
Arista EOS
NVIDIA Spectrum
Juniper JunOS
Prometheus
Grafana
Terraform
REST APIs

Job description

OpenAI is seeking an Infrastructure Operations Engineer to run and improve the large-scale Ethernet fabrics that support GPU clusters, storage, and management networks. You will blend hands-on production operations with automation, observability, and incident response across a global AI network.

The ideal candidate has 5+ years in high-availability data center or cloud network operations and can handle physical-layer troubleshooting as well as routing, fabric behavior, and RCA.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Infrastructure Network Operations Engineer
AI Infrastructure Network Operations Engineer

OpenAI • San Francisco (CA)

On-site
USD 140,000 - 210,000
Network Operations Engineer, AI Networking
Network Operations Engineer, AI Networking

OpenAI • San Francisco (CA)

On-site
USD 140,000 - 210,000
Network Operations Engineer, AI Networking
Network Operations Engineer, AI Networking

Precision Labs • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 210,000
Global HPC Network Engineer for AI Infra
Global HPC Network Engineer for AI Infra

Together • San Francisco (CA)

On-site
USD 190,000 - 280,000
Startup equity
Health insurance
Competitive benefits
Senior Network Engineer — AI Infra & HPC Fabric Expert
Senior Network Engineer — AI Infra & HPC Fabric Expert

Nscale • Houston (TX)

On-site
USD 150,000 - 210,000
Competitive benefits package
Flexible paid time off
Parental leave
+1
AI DataCenter Network Production Engineer
AI DataCenter Network Production Engineer

Intelletec Energy • San Francisco (CA)

On-site
USD 200,000 - 275,000
Senior Network Automation Engineer (Production)
Senior Network Automation Engineer (Production)

Meta • Seattle (WA)

On-site
USD 154,000 - 217,000
Senior Network Engineer — AI Cloud Networking & Automation
Senior Network Engineer — AI Cloud Networking & Automation

Nscale • New York (NY), San Francisco (CA), Seattle (WA)

On-site
USD 140,000 - 190,000
Network Engineer
Network Engineer

OpenAI • San Francisco (CA)

Hybrid
USD 150,000 - 230,000
Relocation assistance
Hybrid work model
Network Production Engineer, Network.AI
Network Production Engineer, Network.AI

Meta • Menlo Park (CA)

Hybrid
USD 210,000 - 320,000