Senior Technical Support Engineer, AI Networking & NOS

The Consensus

United States

On-site

USD 200,000 - 216,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Upscale AI is seeking a Technical Support Engineer with 10+ years of experience to own end-to-end network bring-up and performance for GPU AI clusters during customer PoCs. You’ll lead switching ASIC pipeline tuning, platform readiness, and fabric performance, turning requirements into reproducible configurations that hit utilization and tail-latency targets at scale.

You’ll debug software, develop automation, validate networking platforms, and work closely with customers and engineering teams

Qualifications

  • Data-center networking with hands-on switch ASIC tuning and platform bring-up.
  • Deep buffer/queueing knowledge (headroom, xon/xoff, dynamic thresholds, VOQ vs shared pools).
  • SONiC, Enterprise-OS
  • Optics & PHY: PAM4 signal integrity, FEC modes, DOM/RS-FEC counters, AN/LT quirks, DAC/AOC selection.
  • Tooling: ethtool, devlink, mlnx_qos, perfquery, switch telemetry (INT/sFlow/ERSPAN), gNMI/REST, Prometheus.
  • Benchmarking: perftest, nccl-tests, iPerf3, traffic generators; reading queue stats, ECN marks, and CNP behavior.
  • EVPN/VXLAN leaf‑spine for AI pods; flowlet or latency‑aware hashing.
  • BlueField DPU offloads, GPUDirect RDMA; NIC QoS (DSCP‑to‑TC, PFCx, GEARBOX/FW).
  • Storage fabrics for AI (NFS-RDMA, NVMe-oF) and their QoS interactions.
  • Python/Ansible for templated ASIC profiles; Git workflows for config promotion.

Responsibilities

  • Design, develop, and maintain features and enhancements for the SONiC NOS platform.
  • Debug, troubleshoot, and resolve issues on SONiC platforms.
  • Develop and execute debugging and troubleshoot infrastructure.
  • Collaborate closely with cross-functional teams including hardware engineers and Test teams.
  • Participate in code reviews, architecture discussions, and documentation efforts.
  • Develop support strategies to root-cause Networking ASICs and Networking Systems issues.
  • Develop debugging tools, for Traffic monitoring and Performance measurements.
  • Debug issues across software, Linux systems, networking stacks, and distributed infrastructure.
  • Be a point of contact for customer deployments, integration testing, proof-of-concepts, and field issue resolution.
  • Build tools that improve deployment efficiency, observability, telemetry, and automated testing.
  • Collaborate with software, infrastructure, QA, and product teams to identify root causes and deliver robust solutions.
  • Contribute to backend services, APIs, orchestration components, and infrastructure automation.
  • Document technical findings and communicate effectively with engineering teams and customers.

Skills

Python
Go
Linux
TCP/IP networking
Network troubleshooting

Education

Bachelor’s or Master’s degree in Computer Science or Electrical Engineering

Tools

ethtool
devlink
mlnx_qos
perfquery
Prometheus
gNMI/REST
Docker

Job description

Upscale AI is seeking a Technical Support Engineer with 10+ years of experience to own end-to-end network bring-up and performance for GPU AI clusters during customer PoCs. You’ll lead switching ASIC pipeline tuning, platform readiness, and fabric performance, turning requirements into reproducible configurations that hit utilization and tail-latency targets at scale.

You’ll debug software, develop automation, validate networking platforms, and work closely with customers and engineering teams

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Principal AI Networking Engineer for GPU Clusters
Principal AI Networking Engineer for GPU Clusters

The Consensus • United States

On-site
USD 248,000 - 269,000
Technical Support Senior Staff Engineer
Technical Support Senior Staff Engineer

The Consensus • United States

On-site
USD 200,000 - 216,000
Technical Support Principal Engineer – AI Network
Technical Support Principal Engineer – AI Network

The Consensus • United States

On-site
USD 248,000 - 269,000
Principal AI Networking Engineer — SONiC & 100G+
Principal AI Networking Engineer — SONiC & 100G+

Yoh Services LLC • Santa Clara (CA)

On-site
USD 248,000 - 269,000
Medical, Prescription, Dental & Vision
Health Savings Account (HSA)
Life & Disability Insurance
+4
Senior Network Engineer, AI Infra & High-Performance Cloud
Senior Network Engineer, AI Infra & High-Performance Cloud

Nscale • San Francisco (CA)

On-site
USD 150,000 - 210,000
Medical insurance
Retirement plan
Flexible PTO
Senior AI Networking Solutions Engineer
Senior AI Networking Solutions Engineer

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 168,000 - 322,000
Equity
Benefits package
Senior Network Engineer — AI Infra & HPC Fabric Expert
Senior Network Engineer — AI Infra & HPC Fabric Expert

Nscale • Houston (TX)

On-site
USD 150,000 - 210,000
Competitive benefits package
Flexible paid time off
Parental leave
+1
Senior HPC & AI Networking Engineer (Onsite)
Senior HPC & AI Networking Engineer (Onsite)

NVIDIA Gruppe • Town of Texas (WI)

Hybrid
USD 120,000 - 207,000
Equity options
Comprehensive benefits
Principal AI Networking Support Engineer
Principal AI Networking Support Engineer

Yoh, A Day & Zimmermann Company • California (MO)

On-site
USD 248,000 - 269,000
Medical benefits
Health Savings Account
Life & Disability Insurance
+6
Senior ASIC Modeling Engineer — AI Networking Simulator Lead
Senior ASIC Modeling Engineer — AI Networking Simulator Lead

Doist • United States

On-site
USD 120,000 - 180,000