Senior Network Engineer

Together

San Francisco (CA)

On-site

USD 190,000 - 280,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Startup equity
Health insurance
Competitive benefits

Job summary

Together AI is seeking a Senior Network Engineer to design, deploy, and operate our global, multi-vendor data center networks supporting production AI compute environments.

This hands-on role requires deep networking expertise and comfort troubleshooting across Linux, Kubernetes, automation, and application boundaries to keep networks highly available and performant.

You will shape standards, drive automation, and collaborate with infra, security, and software teams across locations.

Qualifications

  • 8+ years of experience designing, building, and supporting large-scale production networks.
  • Deep TCP/IP knowledge with BGP, OSPF, VXLAN, EVPN, ECMP, QoS.
  • Experience with multi-tenant networks using VRFs, VLANs, overlays, and policy-based segmentation.
  • Hands-on experience with Arista, Cisco, Juniper, and NVIDIA hardware.
  • Strong troubleshooting using Wireshark, tcpdump, MTR, curl, nmap, and Linux utilities.
  • Ability to diagnose connectivity, latency, packet loss, routing, and performance across network, host, and application layers.
  • Experience developing or maintaining network automation using Python, Ansible, or similar.
  • Experience with Git-based SDLC: branching, code review, validation, linting, testing, CI/CD, deployment, rollback.
  • Working knowledge of Kubernetes networking (pods, services, CNIs).
  • Foundational knowledge of RDMA (RoCE/InfiniBand).
  • Experience with cloud networking (AWS, GCP, Azure).
  • Strong Linux administration and troubleshooting.

Responsibilities

  • Design, deploy, operate, and maintain global, multi-vendor networks for high-performance AI compute infra.
  • Troubleshoot complex network and application connectivity issues, identify root causes, drive resolution.
  • Analyze telemetry, packet captures, logs, and performance data to identify degradation and capacity constraints.
  • Participate in architecture reviews to ensure performance, availability, security, and operability.
  • Develop automation, validation, and tooling to improve network reliability and reduce manual effort.
  • Evaluate network hardware, software, optics, and emerging technologies for production use.
  • Establish standards and best practices for network design, deployment, monitoring, change management, and incident response.
  • Lead projects addressing complex technical challenges and contribute to the network engineering roadmap.
  • Collaborate with infra, systems, security, and application teams to troubleshoot cross-boundary issues.

Skills

TCP/IP fundamentals
BGP/OSPF/VXLAN
Multi-tenant networks (VRFs/VLANs)
Vendor hardware: Arista/Cisco/Juniper
Network troubleshooting (Wireshark/tcp

Tools

Wireshark
tcpdump
MTR
curl
nmap
Python
Ansible

Job description

Together AI is looking for a Senior Network Engineer to design, deploy, and operate the global network infrastructure supporting our production services and high-performance AI compute environments.

This is a hands-on engineering role for someone with deep networking expertise who can also troubleshoot across Linux, Kubernetes, automation, and application boundaries. You will work on large-scale, multi-vendor data center networks and help ensure they remain highly available, reliable, scalable, and performant.

The ideal candidate has strong networking fundamentals, experience operating complex networks at scale, and a structured, evidence-based approach to troubleshooting. You should be comfortable owning problems from initial investigation through root cause and resolution, including situations where the issue may extend beyond the network itself.

Requirements
  • 8+ years of professional experience designing, building, and supporting large-scale production data center, cloud, service-provider, or high-performance computing networks (excluding enterprise networks).
  • Deep understanding of TCP/IP and strong experience with technologies such as BGP, OSPF, VXLAN, EVPN, ECMP, and QoS.
  • Experience designing and supporting multi-tenant network environments using technologies such as VRFs, VLANs, overlays, and policy-based segmentation.
  • Hands-on experience deploying and troubleshooting network platforms from vendors such as Arista, Cisco, Juniper, and NVIDIA.
  • Strong troubleshooting skills using tools such as Wireshark, tcpdump, MTR, curl, nmap, and standard Linux networking utilities.
  • Ability to diagnose connectivity, latency, packet-loss, routing, and performance issues across the network, host, and application layers.
  • Experience developing or maintaining network automation using Python, Ansible, or similar tools.
  • Experience working through a Git-based software development lifecycle, including branching, code review, validation, linting, testing, CI/CD, deployment, and rollback.
  • Working knowledge of Kubernetes networking, including pods, services, CNIs, and basic connectivity troubleshooting.
  • Foundational knowledge of RDMA networking and technologies such as RoCE or InfiniBand.
  • Experience with cloud networking in AWS, GCP, or Azure.
  • Strong Linux administration and troubleshooting skills.
Responsibilities
  • Design, deploy, operate, and maintain global, multi-vendor, multi-protocol networks supporting high-performance AI compute infrastructure.
  • Troubleshoot complex network and application-connectivity issues, identify root causes, and drive problems through resolution.
  • Analyze telemetry, packet captures, logs, and performance data to identify network degradation, congestion, packet loss, and capacity constraints.
  • Participate in architecture and design reviews to ensure solutions meet requirements for performance, availability, scalability, security, and operational supportability.
  • Develop and maintain automation, validation, and operational tooling that improves network reliability and reduces manual effort.
  • Evaluate network hardware, software, optics, and emerging technologies for use in production environments.
  • Establish standards and operational best practices for network design, deployment monitoring, change management, and incident response.
  • Lead projects addressing complex technical challenges and contribute directly to the network engineering roadmap.
  • Partner with infrastructure, systems, security, and application teams to troubleshoot issues that cross traditional ownership boundaries.
Preferred
  • Hands-on experience deploying or operating RoCE and/or InfiniBand fabrics.
  • Experience supporting GPU clusters, HPC environments, distributed storage, or other high-bandwidth and latency-sensitive workloads.
  • Understanding of AI training and inference traffic patterns and the demands they place on network infrastructure.
  • Experience operating networks spanning thousands of devices, multiple data centers, and multiple geographic regions.
  • Familiarity with AI-assisted engineering tools and the ability to validate, test, and safely deploy AI-generated automation or code.
About Together AI

Together AI is a research-driven artificial intelligence company. We believe open and transparent AI systems will drive innovation and create the best outcomes for society, and together we are on a mission to significantly lower the cost of modern AI systems by co-designing software, hardware, algorithms, and models. We have contributed to leading open-source research, models, and datasets to advance the frontier of AI, and our team has been behind technological advancement such as FlashAttention, Hyena, FlexGen, and RedPajama. We invite you to join a passionate group of researchers and engineers in our journey in building the next generation AI infrastructure.

Compensation

We offer competitive compensation, startup equity, health insurance and other competitive benefits.

The US base salary range for this full-time position is: $190,000 - $280,000 + equity + benefits. Our salary ranges are determined by location, level and role. Individual compensation will be determined by experience, skills, and job-related knowledge.

Together AI is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Network Engineer
Senior Network Engineer

Togetherai • San Francisco (CA)

On-site
USD 190,000 - 250,000
Health insurance
Startup equity
Competitive benefits
Senior Network Engineer
Senior Network Engineer

Together AI • San Francisco (CA)

On-site
USD 190,000 - 270,000
Equity
Health insurance
Competitive pay
Senior Network Engineer
Senior Network Engineer

AI Chopping Block • San Francisco (CA)

On-site
USD 190,000 - 270,000
Equity
Health insurance
Startup benefits
Senior Software Engineer - Together Cloud Infrastructure
Senior Software Engineer - Together Cloud Infrastructure

Togetherai • San Francisco (CA)

Hybrid
USD 160,000 - 230,000
Equity
Health insurance
Flexible remote work options
AI Infrastructure Systems Engineer
AI Infrastructure Systems Engineer

Together • San Francisco (CA)

On-site
USD 190,000 - 270,000
Health insurance
Startup equity
Competitive benefits
Senior Backend Engineer, Inference Platform
Senior Backend Engineer, Inference Platform

Together • San Francisco (CA)

On-site
USD 160,000 - 250,000
Equity
Health insurance
Competitive compensation
Senior Software Engineer - Together Cloud Platform
Senior Software Engineer - Together Cloud Platform

Togetherai • San Francisco (CA)

Hybrid
USD 160,000 - 230,000
Startup equity
Health insurance
Flexible remote work options
Senior Software Engineer - Together Cloud Infrastructure
Senior Software Engineer - Together Cloud Infrastructure

Together AI • San Francisco (CA)

On-site
USD 160,000 - 230,000
Startup Equity
Health Insurance
Flexible Remote Work
Senior Backend Engineer, Inference Platform
Senior Backend Engineer, Inference Platform

Togetherai • San Francisco (CA)

On-site
USD 160,000 - 250,000
Health insurance
Startup equity
Competitive benefits
Associate, Infrastructure Strategy & Operations
Associate, Infrastructure Strategy & Operations

Together AI • San Francisco (CA)

Remote
USD 140,000 - 170,000
Startup equity
Health insurance
Remote work flexibility
+1