Senior Network & Site Reliability Engineer

Alembic Technologies

San Francisco (CA)

On-site

USD 210,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Alembic Technologies in San Francisco is seeking an experienced engineer to design and operate the global network for a high-performance private supercomputer. You will architect scalable, secure networks and manage incident response systems, ensuring operational readiness.

The ideal candidate has over 8 years in infrastructure engineering, extensive experience with network devices, and strong skills in automation tools. Join us in building resilient systems for machine learning workloads and real-time analytics.

Qualifications

  • 8+ years in network or infrastructure engineering.
  • Experience in datacenter operations and network administration.
  • Strong background in network architecture and operations.

Responsibilities

  • Architect and operate scalable network architecture.
  • Own network device configuration management.
  • Build and maintain monitoring and incident response systems.

Skills

Network security
Infrastructure engineering
Network automation
Scripting (Python, Bash)
Monitoring and observability tools

Tools

Ansible
Terraform
Prometheus
Grafana

Job description

About The Role

We’re building infrastructure that has to perform under real‑world scale, reliability, and security demands – and we’re looking for an engineer who wants to own the foundation it runs on. This isn’t a traditional "keep the lights on" role. You’ll design and operate the global network and reliability layer behind one of the world’s fastest private supercomputers – the fabric powering distributed compute, ML workloads, real‑time analytics, and mission‑critical enterprise systems. You’ll work across networking, systems, automation, observability, and reliability engineering to scale a platform where performance genuinely matters, with real influence over architecture decisions. It’s a strong fit if you like solving deep infrastructure problems, building resilient systems, automating everything repetitive, and owning architecture rather than just maintaining it.

What You'll Do
  • Architect and operate scalable, secure network architecture for high‑security requirements and large‑scale machine learning workloads.
  • Own network device configuration management end to end, ensuring consistency and reliability across the fleet.
  • Improve system and network reliability and performance through automation, observability, and proactive capacity planning.
  • Implement and manage complex network protocols and connectivity, including BGP, VPNs, and WAN circuits and external peering.
  • Build and maintain comprehensive monitoring, alerting, and incident response – SLOs, runbooks, and on‑call rotations – and drive post‑incident analysis and continuous improvement.
  • Ensure security, compliance, and operational readiness across our network and cloud infrastructure.
  • Partner across engineering and data science to drive a culture of performance and reliability.
What Will Help You Succeed
  • 8+ years in network or infrastructure engineering, including 5+ years in datacenter operations and/or systems and network administration.
  • A strong background in network security, architecture, design, and operations.
  • Extensive hands‑on experience with network devices (firewalls, switches, load balancers) and large‑scale architectures and protocols – BGP, QoS, MPLS, and IPsec VPNs.
  • Experience designing and operating modern datacenter network fabrics (spine‑leaf, EVPN/VXLAN, ECMP).
  • Network automation and IaC tooling (Ansible, Terraform, Nornir, or similar), plus IPAM/DCIM platforms (NetBox, Infoblox, or similar).
  • WAN engineering – carrier circuit provisioning and external network peering.
  • Familiarity with Kubernetes networking (CNI plugins, ingress, service networking, network policy) and strong operational experience with Linux‑based production infrastructure.
  • Experience with monitoring and observability stacks (Prometheus, Grafana, Datadog, ELK, OpenTelemetry).
  • Solid scripting (Python, Bash) to debug complex network and system issues and automate solutions, plus excellent cross‑functional communication.
Also Helpful
  • NVIDIA networking technologies – Cumulus Linux, InfiniBand, Spectrum‑X, and BlueField DPUs (this is the fabric behind our SuperPOD).
  • Familiarity with data‑intensive platforms (Spark, Airflow, Kafka) and storage network protocols (NFS, LustreFS, iSCSI).
  • Security practices for applications and infrastructure, and experience in high‑compliance or SOC 2 environments.
The Role Is Right for You If
  • You want to own mission‑critical network and infrastructure end to end – from architecture to incident management – not just keep it running.
  • You’d rather build and automate than direct from a distance, and you want meaningful influence over how a high‑performance platform scales.

Compensation Range: $210K - $240K

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Network Engineer
Senior Network Engineer

Nscale • United States

On-site
USD 150,000 - 210,000
Medical, dental, vision
Flexible paid time off
Parental leave
+1
Senior Network Security Engineer
Senior Network Security Engineer

GTN Technical Staffing • Dallas (TX)

On-site
USD 150,000 - 230,000
Principal Network Engineer - Deployments
Principal Network Engineer - Deployments

Nscale • Seattle (WA)

On-site
USD 220,000 - 280,000
Medical insurance
Dental insurance
Vision insurance
+3
Senior Data Center Network Engineer – AI/HPC Infrastructure
Senior Data Center Network Engineer – AI/HPC Infrastructure

StratITech • San Francisco (CA)

On-site
USD 210,000 - 240,000
Equity
Network Engineer, Supercomputing
Network Engineer, Supercomputing

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 350,000 - 475,000
Health, dental, and vision benefits
Unlimited PTO
Parental leave
+1
Senior Software Engineer, Network Automation
Senior Software Engineer, Network Automation

NMC2 • Dallas (TX)

On-site
USD 90,000 - 130,000
Senior Software Engineer, Network Automation
Senior Software Engineer, Network Automation

NorthMark Compute & Cloud • Dallas (TX)

On-site
USD 90,000 - 120,000
Senior Network Engineer
Senior Network Engineer

Nscale • Seattle (WA)

On-site
USD 150,000 - 210,000
Principal Network Engineer
Principal Network Engineer

nscaleoperationsukltd • Seattle (WA)

On-site
USD 180,000 - 260,000
Network Architect
Network Architect

Tata Consultancy Services • Charlotte (NC)

Hybrid
USD 95,000 - 120,000