Global HPC Network Engineer for AI Infra

Together

San Francisco (CA)

On-site

USD 190,000 - 280,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Startup equity
Health insurance
Competitive benefits

Job summary

Together AI is seeking a Senior Network Engineer to design, deploy, and operate our global, multi-vendor data center networks supporting production AI compute environments.

This hands-on role requires deep networking expertise and comfort troubleshooting across Linux, Kubernetes, automation, and application boundaries to keep networks highly available and performant.

You will shape standards, drive automation, and collaborate with infra, security, and software teams across locations.

Qualifications

  • 8+ years of experience designing, building, and supporting large-scale production networks.
  • Deep TCP/IP knowledge with BGP, OSPF, VXLAN, EVPN, ECMP, QoS.
  • Experience with multi-tenant networks using VRFs, VLANs, overlays, and policy-based segmentation.
  • Hands-on experience with Arista, Cisco, Juniper, and NVIDIA hardware.
  • Strong troubleshooting using Wireshark, tcpdump, MTR, curl, nmap, and Linux utilities.
  • Ability to diagnose connectivity, latency, packet loss, routing, and performance across network, host, and application layers.
  • Experience developing or maintaining network automation using Python, Ansible, or similar.
  • Experience with Git-based SDLC: branching, code review, validation, linting, testing, CI/CD, deployment, rollback.
  • Working knowledge of Kubernetes networking (pods, services, CNIs).
  • Foundational knowledge of RDMA (RoCE/InfiniBand).
  • Experience with cloud networking (AWS, GCP, Azure).
  • Strong Linux administration and troubleshooting.

Responsibilities

  • Design, deploy, operate, and maintain global, multi-vendor networks for high-performance AI compute infra.
  • Troubleshoot complex network and application connectivity issues, identify root causes, drive resolution.
  • Analyze telemetry, packet captures, logs, and performance data to identify degradation and capacity constraints.
  • Participate in architecture reviews to ensure performance, availability, security, and operability.
  • Develop automation, validation, and tooling to improve network reliability and reduce manual effort.
  • Evaluate network hardware, software, optics, and emerging technologies for production use.
  • Establish standards and best practices for network design, deployment, monitoring, change management, and incident response.
  • Lead projects addressing complex technical challenges and contribute to the network engineering roadmap.
  • Collaborate with infra, systems, security, and application teams to troubleshoot cross-boundary issues.

Skills

TCP/IP fundamentals
BGP/OSPF/VXLAN
Multi-tenant networks (VRFs/VLANs)
Vendor hardware: Arista/Cisco/Juniper
Network troubleshooting (Wireshark/tcp

Tools

Wireshark
tcpdump
MTR
curl
nmap
Python
Ansible

Job description

Together AI is seeking a Senior Network Engineer to design, deploy, and operate our global, multi-vendor data center networks supporting production AI compute environments.

This hands-on role requires deep networking expertise and comfort troubleshooting across Linux, Kubernetes, automation, and application boundaries to keep networks highly available and performant.

You will shape standards, drive automation, and collaborate with infra, security, and software teams across locations.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Network Architect — Global HPC Data Center Networks
Senior Network Architect — Global HPC Data Center Networks

Together AI • San Francisco (CA)

On-site
USD 190,000 - 270,000
Equity
Health insurance
Competitive pay
Senior Network Engineer — AI Infra & HPC Fabric Expert
Senior Network Engineer — AI Infra & HPC Fabric Expert

Nscale • Houston (TX)

On-site
USD 150,000 - 210,000
Competitive benefits package
Flexible paid time off
Parental leave
+1
Staff-Level Data Center Network Engineer, AI/HPC Infra
Staff-Level Data Center Network Engineer, AI/HPC Infra

StratITech • San Francisco (CA)

On-site
USD 210,000 - 240,000
Equity
Network Architect
Network Architect

BayOne Solutions • Bellevue (WA)

On-site
USD 150,000 - 190,000
Senior HPC-AI Cluster Architect (Equity)
Senior HPC-AI Cluster Architect (Equity)

NVIDIA • Santa Clara (CA)

On-site
USD 176,000 - 334,000
Equity
Benefits
Senior AI/HPC Networking Solutions Architect
Senior AI/HPC Networking Solutions Architect

NVIDIA AI • Indiana (PA)

On-site
USD 140,000 - 210,000
Senior Network Engineer — AI Cloud Networking & Automation
Senior Network Engineer — AI Cloud Networking & Automation

Nscale • New York (NY), San Francisco (CA), Seattle (WA)

On-site
USD 140,000 - 190,000
Remote AI Infra & HPC Pre-Sales Architect
Remote AI Infra & HPC Pre-Sales Architect

Levelociti • United States

Remote
USD 140,000 - 190,000
Senior Network Engineer - AI Cloud Infra & InfiniBand
Senior Network Engineer - AI Cloud Infra & InfiniBand

Cirrascale • Town of Texas (WI), Northern (KY)

Hybrid
USD 125,000 - 180,000
Health, dental, vision insurance
Paid time off
Professional development
Senior AI/HPC Solutions Architect – Linux & Networking
Senior AI/HPC Solutions Architect – Linux & Networking

Nvidia Corporation • Santa Clara (CA)

On-site
USD 148,000 - 236,000