Principal AI Networking Engineer for GPU Clusters

The Consensus

United States

On-site

USD 248,000 - 269,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Upscale AI is seeking a Technical Support Lead Engineer with 12+ years of experience to own end-to-end network bring-up for GPU AI clusters and PoCs. You will lead ASIC pipeline tuning, platform readiness, and fabric performance, turning requirements into reproducible configurations that meet utilization and tail-latency targets at scale.

Join the AI Networking team to debug software, develop automation, validate platforms, and work with customers and engineers to ensure outstanding deployment

Qualifications

  • Bachelor’s or master’s degree in computer science, Electrical Engineering, or a related field.
  • Minimum of 12 years of work experience is required, with at least 3 years of hands-on SONiC or equivalent Network Operating System (NOS) development experience preferred.
  • Strong programming skills in Python, Go, or a similar language.
  • Solid understanding of Linux, TCP/IP networking, routing, switching, VLANs, and network troubleshooting.
  • Experience with PTF (Packet Test Framework) and SPyTest for network validation.
  • Familiarity with Linux internals, docker containers.
  • Excellent problem-solving skills and ability to work in a fast-paced, collaborative environment.
  • Knowledge of network ASICs and switch hardware architecture is mandatory.
  • Excellent written and verbal communication skills.
  • Ability to thrive in a collaborative, fast-paced startup environment with a strong sense of ownership.

Responsibilities

  • Design, develop, and maintain features and enhancements for the SONiC NOS platform.
  • Debug, troubleshoot, and resolve issues on SONiC platforms.
  • Develop and execute debugging and troubleshoot infrastructure.
  • Collaborate closely with cross-functional teams including hardware engineers and Test teams.
  • Participate in code reviews, architecture discussions, and documentation efforts.
  • Develop support strategies to root-cause Networking ASICs and Networking Systems issues.
  • Develop debugging tools, for Traffic monitoring and Performance measurements.
  • Debug issues across software, Linux systems, networking stacks, and distributed infrastructure.
  • Be a point of contact for customer deployments, integration testing, proof-of-concepts, and field issue resolution.
  • Build tools that improve deployment efficiency, observability, telemetry, and automated testing.
  • Collaborate with software, infrastructure, QA, and product teams to identify root causes and deliver robust solutions.
  • Contribute to backend services, APIs, orchestration components, and infrastructure automation.
  • Document technical findings and communicate effectively with engineering teams and customers.

Skills

Python
Go
Linux
TCP/IP networking
Network troubleshooting

Education

Bachelor's or master's degree in CS/EE
12+ years experience in networking/SONiC NOS

Tools

PTF
SPyTest
Docker
ethtool
Devlink

Job description

Upscale AI is seeking a Technical Support Lead Engineer with 12+ years of experience to own end-to-end network bring-up for GPU AI clusters and PoCs. You will lead ASIC pipeline tuning, platform readiness, and fabric performance, turning requirements into reproducible configurations that meet utilization and tail-latency targets at scale.

Join the AI Networking team to debug software, develop automation, validate platforms, and work with customers and engineers to ensure outstanding deployment

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead AI Networking & Infrastructure Engineer
Lead AI Networking & Infrastructure Engineer

Upscale AI • United States

On-site
USD 248,000 - 269,000
Senior Technical Support Engineer, AI Networking & NOS
Senior Technical Support Engineer, AI Networking & NOS

The Consensus • United States

On-site
USD 200,000 - 216,000
Senior AI Networking Engineer – GPU Cluster PoC Lead
Senior AI Networking Engineer – GPU Cluster PoC Lead

Upscale AI • United States

On-site
USD 200,000 - 216,000
Senior HPC-AI Cluster Architect (Equity)
Senior HPC-AI Cluster Architect (Equity)

NVIDIA • Santa Clara (CA)

On-site
USD 176,000 - 334,000
Equity
Benefits
Senior AI Network Architect for Large-Scale GPU Clusters
Senior AI Network Architect for Large-Scale GPU Clusters

Hamilton Barnes Associates Limited • United States

On-site
USD 220,000 - 350,000
Annual bonus
Equity opportunities
Flexible working arrangements
+1
Senior AI Networking & Performance Engineer
Senior AI Networking & Performance Engineer

NVIDIA • Town of Texas (WI)

On-site
USD 272,000 - 432,000
Senior AI/HPC Networking Solutions Architect
Senior AI/HPC Networking Solutions Architect

NVIDIA AI • Indiana (PA)

On-site
USD 140,000 - 210,000
Lead AI Networking Engineer - Equity
Lead AI Networking Engineer - Equity

NVIDIA AI • Seattle (WA)

On-site
USD 180,000 - 240,000
Equity
Health Insurance
AI Grid Networking Engineer - GPU Fabric Performance
AI Grid Networking Engineer - GPU Fabric Performance

AMP PBC • San Francisco (CA)

On-site
USD 180,000 - 240,000
Senior Solutions Architect - AI Cluster Networking Design
Senior Solutions Architect - AI Cluster Networking Design

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Benefits