Senior Principal Network Engineer

Graphcore

Austin (TX)

Presencial

USD 130 000 - 160 000

Tempo integral

14 dias+
Gerador de candidaturas

Recebe uma resposta deste empregador — um currículo e uma carta de apresentação adaptados exatamente ao que estão a contratar.

Ultrapassa os filtros ATS

Resumo da oferta

A leading AI compute company based in Austin, Texas is seeking a Senior Principal Network Engineer to design, deploy, and optimize advanced AI data center networks. You'll ensure high-performance computing infrastructure for next-generation workloads. Ideal candidates will have significant experience with network engineering in high-density environments, expertise in data center networking protocols, and proficiency in automation technologies. A collaborative spirit and strong communication skills are essential for this pivotal role.

Qualificações

  • 12+ years of progressive network engineering experience.
  • At least 3 years in hyperscale, high-density, or HPC data center environments.
  • Experience deploying high-speed network technologies including 400G/800G optics.

Responsabilidades

  • Design ultra-high-bandwidth AI network fabrics for distributed AI workloads.
  • Optimize performance of lossless Ethernet fabrics using congestion control mechanisms.
  • Lead initiatives to implement NetDevOps practices and develop automation.

Conhecimentos

Expert-level knowledge of data center routing and switching protocols
Strong operational understanding of RDMA networking technologies
Proficiency in automation and scripting languages
Experience operating large-scale AI or GPU clusters
Strong collaboration and communication skills

Formação académica

BS or MS in Computer Science, Electrical Engineering, Network Engineering

Ferramentas

Python
Go
Bash

Descrição da oferta de emprego

Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry.

As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone.

Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation.

Job Summary

We are seeking a Senior Principal Network Engineer to help design, deploy, and optimize next‑generation AI data center networks. AI training and inference workloads require extremely high bandwidth, deterministic low latency, and zero‑packet‑loss networking environments. In this role, you will partner closely with the Network Architecture Lead to design and scale high‑performance computing (HPC) network fabrics supporting GPU clusters. You will work across hardware, networking, and AI application layers to ensure Graphcore’s large‑scale AI infrastructure operates at peak performance. The ideal candidate brings deep experience operating hyperscale or HPC data center networks and has expertise in high‑speed Ethernet fabrics, RDMA technologies, advanced automation, and telemetry systems.

The Team

The Data Center Network Engineering team designs and operates the high‑performance network fabrics that power Graphcore’s AI compute platforms. The team collaborates closely with hardware engineering, AI researchers, and infrastructure teams to build scalable networking environments optimized for distributed training and inference workloads. Engineers work on pioneering technologies including high‑speed Ethernet fabrics, lossless networking, RDMA transport, and large‑scale automation frameworks to support next‑generation AI clusters.

Responsibilities and Duties
  • Assist in defining ultra‑high‑bandwidth, non‑blocking AI network fabrics (Clos spine‑leaf‑super‑spine architectures) for large‑scale distributed AI workloads.
  • Optimize performance of lossless Ethernet fabrics using congestion control mechanisms such as PFC, ECN, and DCQCN to support RDMA/RoCEv2 communication.
  • Lead initiatives to implement NetDevOps practices and develop automation for provisioning, configuration management, and network remediation.
  • Design and deploy high‑resolution telemetry pipelines to monitor network health, detect microbursts, and analyze congestion patterns.
  • Support modeling, deployment, configuration, and monitoring of data center network fabrics including scale‑out, scale‑up, and front‑end networks.
  • Collaborate cross‑functionally with hardware engineers, AI researchers, and data center operations teams to co‑design high‑performance infrastructure.
  • Provide technical leadership and mentorship to network engineers while establishing best practices and operational standards.
  • Contribute to the long‑term networking strategy and roadmap for Graphcore’s AI infrastructure.
  • Research and evaluate next‑generation high‑speed networking technologies and vendor solutions.
Candidate Profile
  • BS or MS or equivalent experience in Computer Science, Electrical Engineering, Network Engineering, or related technical discipline.
  • 12+ years of progressive network engineering experience with at least 3 years in hyperscale, high‑density, or HPC data center environments.
  • Expert‑level knowledge of data center routing and switching protocols including BGP, OSPF, and EVPN‑VXLAN architectures.
  • Strong operational understanding of RDMA networking technologies such as RoCEv2 or InfiniBand.
  • Hands‑on experience with modern merchant silicon networking platforms and NOS platforms such as Arista EOS, Cisco NX‑OS, or SONiC.
  • Experience deploying high‑speed network technologies including 400G/800G optics and large‑scale fabric architectures.
  • Proficiency in automation and scripting languages such as Python, Go, Bash, or similar tools.
  • Strong collaboration and communication skills across cross‑functional engineering teams.
  • Experience operating large‑scale AI or GPU clusters.
  • Familiarity with network telemetry frameworks and streaming analytics.
  • Experience implementing NetDevOps workflows and infrastructure automation pipelines.
  • Experience influencing vendor roadmaps or evaluating next‑generation networking technologies.

As set forth in Graphcore’s Equal Employment Opportunity policy, we do not discriminate on the basis of any protected group status under any applicable law.

Obtém a tua avaliação gratuita e confidencial do currículo.

ou arrasta e larga o ficheiro aqui.

Similar jobs

Ofertas semelhantes que vale a pena comparar

Senior Principal Network Engineer
Senior Principal Network Engineer

Graphcore • Austin (TX)

Presencial
USD 140 000 - 180 000
Principal Network Engineer
Principal Network Engineer

Graphcore • Milpitas (CA)

Presencial
USD 120 000 - 160 000
Senior AI Data Center Network Engineer & Automation
Senior AI Data Center Network Engineer & Automation

Graphcore • Milpitas (CA)

Presencial
USD 120 000 - 160 000
Principal Engineer
Principal Engineer

Graphcore • Milpitas (CA), Austin (TX)

Presencial
USD 180 000 - 260 000
Medical, dental, and vision
401(k) retirement plan
Commuter benefits
+2
Principal Datacenter Technologist Multiple Vacancies Available
Principal Datacenter Technologist Multiple Vacancies Available

Graphcore • Milpitas (CA), Northern (KY)

Híbrido
USD 200 000 - 280 000
Medical, dental, and vision coverage
Flexible working hours
Hybrid working arrangements
+1
Principal Datacenter Technologist
Principal Datacenter Technologist

Graphcore • Milpitas (CA)

Híbrido
USD 180 000 - 240 000
Hybrid working arrangements
Professional development resources
Mentorship opportunities
Principal Engineer
Principal Engineer

Graphcore • EUA

Presencial
USD 180 000 - 240 000
Medical coverage
Dental coverage
Vision coverage
+4
Staff AI Performance Engineer
Staff AI Performance Engineer

EngineersOfAI • Austin (TX)

Presencial
USD 90 000 - 120 000
GPU Network Engineer
GPU Network Engineer

Blue Signal Search • Santa Clara (CA)

Presencial
USD 180 000 - 240 000
Technical Services Director, Global Data Center & Lab Infrastructure
Technical Services Director, Global Data Center & Lab Infrastructure

Graphcore • Austin (TX)

Presencial
USD 180 000 - 240 000
401(k) matching
Flexible PTO
Professional development