Network Reliability Engineer — Automation & AI Tooling

Fluidstack

Austin (TX)

On-site

USD 208,000 - 269,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Salary growth potential
Premium health benefits

Job summary

Fluidstack is building a global datacenter network and seeks a senior network engineer to own fleet health, automation, and reliability at scale. You will define monitoring, build dashboards, and develop tooling to diagnose, repair, and validate new sites and hardware across multiple datacenters.

You will lead the automation of network operations, build end-to-end monitoring and incident response, and collaborate across teams to deliver resilient network infrastructure.

Qualifications

  • Toil as a bug: build the tool that does it for you.
  • Think in systems: understand how faults propagate and be able to tell them apart.
  • Move toward ambiguity, not away from it: map the fog and explain it.
  • Learn at a steep slope: reach real competence quickly in unfamiliar domains.
  • Carry a pager: run incidents, write postmortems, fix root causes.
  • Fluent with AI tooling: LLM APIs, MCP servers, and agentic frameworks used daily.
  • Shipped production network tooling or automation relied on by other teams; comfortable in any language with AI coding tools.
  • Developed automation tools in Go and Python; experienced in link diagnostics, optical networks, and network monitoring (gNMI, gRPC, NETCONF, SONiC).
  • Bonus: RMA and repair lifecycle automation; large-scale datacenter fabric (BGP, ECMP, spine-leaf); out-of-band management.

Responsibilities

  • Own network fleet health end to end. Define realtime monitoring requirements, build the alerting lifecycle, and ship dashboards that give on-call engineers a true picture of network state across all sites.
  • Build active debugging tooling. Link diagnostics, remote command execution across the fleet, and repair visualization to turn faults into solvable problems quickly.
  • Turn repair into a pipeline: automate detection through parts management and return to service; include ticket integration, repair lifecycle pipelines, and transceiver/optic tracking.
  • Own network qualification and validation. Build frameworks that gate new sites and hardware into production; define what healthy network looks like before traffic
  • Own end-to-end reliability, scalability, and operation of the network at-scale; Fluidstack aims to build one of the world's largest datacenter networks with aggressive automation and incident discipline.

Skills

Go
Python
Network automation
gNMI
gRPC
NETCONF
AI tooling
SSH

Job description

Fluidstack is building a global datacenter network and seeks a senior network engineer to own fleet health, automation, and reliability at scale. You will define monitoring, build dashboards, and develop tooling to diagnose, repair, and validate new sites and hardware across multiple datacenters.

You will lead the automation of network operations, build end-to-end monitoring and incident response, and collaborate across teams to deliver resilient network infrastructure.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Network Production Engineer - Reliability & Automation (Equity)
Network Production Engineer - Reliability & Automation (Equity)

Fluidstack • New York (NY)

On-site
USD 173,000 - 224,000
Health insurance
Equity
Retirement plan
+1
Data Center IT Engineer - Automation & Multi-Site Ops
Data Center IT Engineer - Automation & Multi-Site Ops

Fluidstack • Austin (TX)

On-site
USD 218,000 - 252,000
Senior Data Center Reliability Engineer
Senior Data Center Reliability Engineer

Fluidstack • Austin (TX)

On-site
USD 220,000 - 260,000
Network Deployment Architect Lead — Scale AI Fabric
Network Deployment Architect Lead — Scale AI Fabric

Fluidstack • San Francisco (CA)

On-site
USD 186,000 - 220,000
Reliability Engineer - Large-Scale AI & HPC Systems
Reliability Engineer - Large-Scale AI & HPC Systems

PVH (Tommy Hilfiger/Calvin Klein) • San Francisco (CA)

On-site
USD 120,000 - 180,000
Staff Network Engineer: Scale AI Infra with Automation
Staff Network Engineer: Scale AI Infra with Automation

Fluidstack • Seattle (WA), New York (NY), San Francisco (CA), Austin (TX)

On-site
USD 180,000 - 240,000
Principal Data Center Reliability Engineer
Principal Data Center Reliability Engineer

Fluidstack • United States

On-site
USD 220,000 - 260,000
Production Network Engineer: Automate & Repair
Production Network Engineer: Automate & Repair

Socket.dev • San Francisco (CA)

On-site
USD 140,000 - 190,000
Senior Network Engineering Lead for AI Compute Scale
Senior Network Engineering Lead for AI Compute Scale

Fluidstack • Seattle (WA), San Francisco (CA), Austin (TX), New York (NY)

On-site
USD 180,000 - 270,000
Senior On-Site Network & Cabling Engineer
Senior On-Site Network & Cabling Engineer

FluidStack • New Lebanon (IN)

On-site
USD 140,000 - 200,000
Equity
Health insurance
Pension plan
+1