AI DataCenter Network Production Engineer

Intelletec Energy

San Francisco (CA)

On-site

USD 200,000 - 275,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Intelletec Energy is a fast-growing hyperscale data center operator building infrastructure to support AI labs and enterprises. We seek a DataCenter Network Production Engineer to own incidents, develop tooling, and drive automated recovery across our fabric and edge deployments.

You will write Python and Go to automate diagnostics, remote commands, and fault-tolerance workflows, while expanding the monitoring and health-visibility layer to support a rapidly expanding fleet.

Qualifications

  • Experience building tooling to diagnose network incidents at scale.
  • Ability to create remote command execution and visualization for failures.

Responsibilities

  • Own incidents in the production network environment through resolution, repair/replacement, and Postmortems.
  • Develop software in Python and Go to automate diagnostics and remote execution.
  • Build and improve the equipment-recovery lifecycle with automated fault detection and ticketing.
  • Operate and evolve monitoring and alerting systems for real-time infrastructure health.
  • Bring new facilities and equipment into production with validation procedures.

Skills

Python
Go

Job description

Our client is a fast-growing hyperscale Data Center Operator building infrastructure to support top AI labs and enterprises.

DataCenter Network Production Engineer
What You’ll Help Build
  • Make network incidents diagnosable at scale. Create tooling that connects telemetry and diagnostics across switches, routers, NICs, and host-to-network paths. Enable engineers to run commands remotely across the infrastructure and quickly visualize failures, dependencies, and likely causes.
  • Create an automated hardware recovery workflow. Build the systems that take an issue from initial detection through troubleshooting, ticket creation, hardware replacement, and restoration of service. The goal is to standardize recovery across data center fabrics, edge infrastructure, and server connectivity without relying on manual coordination.
  • Develop the operational visibility layer for a rapidly expanding fleet. Establish the monitoring, alerting, and health-management infrastructure needed to operate multiple large-scale facilities today and a dramatically larger footprint tomorrow. You'll be building core capabilities that don't yet exist.
What You’ll Own
  • Own incidents in the production network environment through resolution—from identifying the underlying failure to implementing the repair or replacement and confirming the connection is back in service, then Postmortems
  • Develop software in Python and Go to eliminate repetitive troubleshooting work. This includes automated connectivity checks, fleet-wide remote execution, diagnostic workflows, and tools that make complex network failures easier to understand.
  • Build and improve the complete equipment-recovery lifecycle, connecting automated fault detection with service tickets, replacement workflows, component tracking, and final restoration.
  • Operate and evolve the monitoring and alerting systems that provide engineers with an accurate, real-time view of infrastructure health across every facility.
  • Bring new facilities and network equipment into production by developing and executing validation procedures that establish reliability before systems begin carrying live workloads.
Salary & Benefits:
  • $200-275k base salary + meaningful equity.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Data Center Network Reliability Engineer
AI Data Center Network Reliability Engineer

Intelletec Energy • San Francisco (CA)

On-site
USD 200,000 - 275,000
HPC Data Center Production Engineer (Automation) - Banking & Finance
HPC Data Center Production Engineer (Automation) - Banking & Finance

Hamilton Barnes • Illinois

On-site
USD 175,000 - 235,000
Premium health, dental, and vision
Performance-based bonus
Meals and high-end office amenities
+2
Principal Data Center Infrastructure Software Engineer
Principal Data Center Infrastructure Software Engineer

Designworks Talent • Bellevue (WA)

Hybrid
USD 150,000 - 210,000
Lead Network Engineer
Lead Network Engineer

Insight Global • New York (NY)

On-site
USD 120,000 - 150,000
Network Operations Engineer, AI Networking
Network Operations Engineer, AI Networking

OpenAI • San Francisco (CA)

On-site
USD 140,000 - 210,000
AI Network Operations Engineer (Production & Automation)
AI Network Operations Engineer (Production & Automation)

Precision Labs • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 210,000
Member of Technical Staff (Network Infrastructure)
Member of Technical Staff (Network Infrastructure)

Hamilton Barnes Associates Limited • United States

On-site
USD 213,000 - 288,000
Data Center Engineer (Fleet Operations) - AI Infrastructure
Data Center Engineer (Fleet Operations) - AI Infrastructure

Hamilton Barnes Associates Limited • United States

On-site
USD 120,000 - 180,000
Full Benefits
Network Operations Center Engineer | Upto $70/hr
Network Operations Center Engineer | Upto $70/hr

Obsidian • San Francisco (CA)

On-site
USD 120,000 - 160,000
Datacenter Network Architect
Datacenter Network Architect

BayOne Solutions • Bellevue (WA)

On-site
USD 150,000 - 230,000