Senior AI Network Reliability Engineer - Automation-Driven

Crusoe

San Francisco (CA)

On-site

USD 165,000 - 200,000

Full time

9 days ago
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Competitive compensation
Restricted Stock Units
Paid time off & holidays
Health, dental & vision insurance
HSA contributions
Parental leave
Life insurance & disability
Tuition reimbursement
Mental health support
Commuter benefits
Cell phone stipend
401(k) with company match
Volunteer time off

Job summary

Crusoe Cloud seeks a Senior Network Production Operations Engineer to support reliability across global edge, backbone, data center fabrics, and GPU interconnects. This hands-on role focuses on incident response, root cause analysis, and automation to scale AI infrastructure reliability.

The ideal candidate has 5+ years in large-scale network operations, strong Python skills, and experience with BGP/EVPN, RDMA, and multi-vendor gear (Arista, Juniper). Hybrid work model preferred.

Qualifications

  • 5+ years of production network engineering experience in large-scale environments.
  • Strong Python and scripting proficiency for diagnostic tooling and automation.
  • Hands-on experience with observability tools including Grafana, Prometheus, Kentik, ThousandEyes.
  • Experience operating RDMA/RoCE lossless fabrics for GPU or HPC workloads.
  • Solid knowledge of BGP, EVPN-VXLAN, IS-IS, OSPF, MPLS, QoS, and TCP/IP in data centers.
  • Proficiency with Arista EOS and Juniper JunOS in multi-vendor leaf-spine architectures.
  • Comfort operating large device fleets across multi-region environments with on-call responsibilities.
  • Bachelor's degree in Computer Science, Electrical Engineering, or related field, or equivalent experience.

Responsibilities

  • Production Reliability: support uptime across Crusoe's global edge, backbone, data center, and GPU cluster networks.
  • Operational Automation: develop Python tooling to reduce toil and automate remediation workflows.
  • Incident Response: participate in high-severity network events and document postmortems.
  • Root Cause Analysis: perform RCAs using tooling and drive remediation to closure.
  • Observability: build automation on top of monitoring stacks to surface signals faster.
  • Operational Standards: maintain runbooks and SOPs, push for more automation.
  • SLI/SLO Execution: help build dashboards and alerts with Architecture and SRE teams.
  • Collaboration: work with Architecture and SRE to improve day-to-day practices.

Skills

Python scripting
Observability tooling
Incident response
Networking fundamentals
RDMA/RoCE
BGP
Arista EOS
Juniper JunOS

Education

Bachelor's degree in CS/EE

Tools

Grafana
Prometheus
Kentik
ThousandEyes
SNMP
NetFlow

Job description

Crusoe Cloud seeks a Senior Network Production Operations Engineer to support reliability across global edge, backbone, data center fabrics, and GPU interconnects. This hands-on role focuses on incident response, root cause analysis, and automation to scale AI infrastructure reliability.

The ideal candidate has 5+ years in large-scale network operations, strong Python skills, and experience with BGP/EVPN, RDMA, and multi-vendor gear (Arista, Juniper). Hybrid work model preferred.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Network Reliability Engineer — Production & Automation
Senior Network Reliability Engineer — Production & Automation

ProducePay • San Francisco (CA)

On-site
USD 165,000 - 200,000
Competitive compensation
RSUs
Paid time off
+3
Senior AI Networking & SDN Engineer
Senior AI Networking & SDN Engineer

crusoe • Sunnyvale (CA)

On-site
USD 170,000 - 205,000
Health insurance
RSUs
Parental Leave
+4
Senior AI Networking Systems Engineer (SDN)
Senior AI Networking Systems Engineer (SDN)

Crusoe • San Francisco (CA)

On-site
USD 170,000 - 205,000
RSUs
Health insurance
401(k) with match
+4
Senior AI Cloud Platform Engineer - Reliability & Scale
Senior AI Cloud Platform Engineer - Reliability & Scale

Crusoe • San Francisco (CA)

On-site
USD 170,000 - 205,000
Health insurance
RSUs
401(k) match
+2
Global Network Deployment Engineer
Global Network Deployment Engineer

Crusoe • Bellevue (WA)

Hybrid
USD 165,000 - 200,000
Competitive compensation
Equity
PTO & holidays
+7
Senior Network Engineer - AI Infra & Global Networks
Senior Network Engineer - AI Infra & Global Networks

Precision Labs • Dallas (TX), Northern (KY)

Hybrid
USD 125,000 - 148,000
Competitive compensation
Equity packages
Paid time off & holidays
+3
AI Network Operations Engineer (Production & Automation)
AI Network Operations Engineer (Production & Automation)

Precision Labs • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 210,000
Cloud Network Engineer I — Build Scalable AI Infrastructure
Cloud Network Engineer I — Build Scalable AI Infrastructure

US Health Partners, LLC • San Francisco (CA)

On-site
USD 117,000 - 138,000
Equity
PTO
Health insurance
+6
Senior Network Production Operations Engineer
Senior Network Production Operations Engineer

Crusoe • San Francisco (CA)

On-site
USD 165,000 - 200,000
Competitive compensation
Restricted Stock Units
Paid time off & holidays
+10
Senior Network Production Engineer - Global DC & HPC
Senior Network Production Engineer - Global DC & HPC

AI Chopping Block, Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 165,000 - 200,000
Competitive compensation
Equity
Paid time off