Lead AI Networking & Infrastructure Engineer

Upscale AI

United States

On-site

USD 248,000 - 269,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Upscale AI in the United States is seeking a Technical Support Lead Engineer with 12+ years to own end-to-end network bring-up and performance for GPU AI clusters, leading ASIC pipeline tuning and fabric performance. You will design, validate, and optimize networking for scalable AI deployments.

You will debug software, develop automation, validate platforms, and collaborate with hardware and software teams to ensure best customer experiences and robust solutions in a fast-paced startup.

Qualifications

  • Bachelor’s or master’s degree in computer science, Electrical Engineering, or a related field.
  • Minimum of 12 years of work experience is required, with at least 3 years of hands-on SONiC or equivalent Network Operating System (NOS) development experience preferred.
  • Strong programming skills in Python, Go, or a similar language.
  • Solid understanding of Linux, TCP/IP networking, routing, switching, VLANs, and network troubleshooting.
  • Experience with PTF (Packet Test Framework) and SPyTest for network validation.
  • Familiarity with Linux internals, docker containers.
  • Excellent problem-solving skills and ability to work in a fast-paced, collaborative environment.
  • Knowledge of network ASICs and switch hardware architecture is mandatory.
  • Excellent written and verbal communication skills.
  • Ability to thrive in a collaborative, fast-paced startup environment with a strong sense of ownership.

Responsibilities

  • Design, develop, and maintain features and enhancements for the SONiC NOS platform.
  • Debug, troubleshoot, and resolve issues on SONiC platforms.
  • Develop and execute debugging and troubleshoot infrastructure.
  • Collaborate closely with cross-functional teams including hardware engineers and Test teams.
  • Participate in code reviews, architecture discussions, and documentation efforts.
  • Develop support strategies to root-cause Networking ASICs and Networking Systems issues.
  • Develop debugging tools, for Traffic monitoring and Performance measurements.
  • Debug issues across software, Linux systems, networking stacks, and distributed infrastructure.
  • Be a point of contact for customer deployments, integration testing, proof-of-concepts, and field issue resolution.
  • Build tools that improve deployment efficiency, observability, telemetry, and automated testing.
  • Collaborate with software, infrastructure, QA, and product teams to identify root causes and deliver robust solutions.
  • Contribute to backend services, APIs, orchestration components, and infrastructure automation.
  • Document technical findings and communicate effectively with engineering teams and customers.

Skills

Python
Go
Linux networking
Network troubleshooting
Docker
SPyTest
PTF

Education

Bachelor’s or Master’s degree in CS/EE

Tools

PTF
SPyTest
Docker
Linux internals

Job description

Upscale AI in the United States is seeking a Technical Support Lead Engineer with 12+ years to own end-to-end network bring-up and performance for GPU AI clusters, leading ASIC pipeline tuning and fabric performance. You will design, validate, and optimize networking for scalable AI deployments.

You will debug software, develop automation, validate platforms, and collaborate with hardware and software teams to ensure best customer experiences and robust solutions in a fast-paced startup.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Principal AI Networking Engineer for GPU Clusters
Principal AI Networking Engineer for GPU Clusters

The Consensus • United States

On-site
USD 248,000 - 269,000
Senior Technical Support Engineer, AI Networking & NOS
Senior Technical Support Engineer, AI Networking & NOS

The Consensus • United States

On-site
USD 200,000 - 216,000
Senior AI Networking Engineer – GPU Cluster PoC Lead
Senior AI Networking Engineer – GPU Cluster PoC Lead

Upscale AI • United States

On-site
USD 200,000 - 216,000
Lead AI Networking Engineer - Equity
Lead AI Networking Engineer - Equity

NVIDIA AI • Seattle (WA)

On-site
USD 180,000 - 240,000
Equity
Health Insurance
Senior Network Engineer, AI Infra & High-Performance Cloud
Senior Network Engineer, AI Infra & High-Performance Cloud

Nscale • San Francisco (CA)

On-site
USD 150,000 - 210,000
Medical insurance
Retirement plan
Flexible PTO
Senior AI Networking & Performance Engineer
Senior AI Networking & Performance Engineer

NVIDIA • Town of Texas (WI)

On-site
USD 272,000 - 432,000
Senior AI Networking Engineer - High-Performance Systems
Senior AI Networking Engineer - High-Performance Systems

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior Network Engineer — AI Infra & HPC Fabric Expert
Senior Network Engineer — AI Infra & HPC Fabric Expert

Nscale • Houston (TX)

On-site
USD 150,000 - 210,000
Competitive benefits package
Flexible paid time off
Parental leave
+1
Senior ASIC Modeling Engineer — AI Networking Simulator Lead
Senior ASIC Modeling Engineer — AI Networking Simulator Lead

Doist • United States

On-site
USD 120,000 - 180,000
Technical Support Principal Engineer – AI Network
Technical Support Principal Engineer – AI Network

The Consensus • United States

On-site
USD 248,000 - 269,000