Machine Learning Systems Engineer, Networking

NVIDIA Corporation

Santa Clara, Northern (CA, KY)

Hybrid

USD 152,000 - 288,000

Full time

10 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

NVIDIA Corporation is seeking a Machine Learning Systems Engineer, Networking to design and implement ML algorithms for real-time telemetry pipelines across GPU fleets. You will build end-to-end ML pipelines and optimize performance within strict CPU/memory budgets, enabling proactive AI training and inference across massive-scale infrastructure.

The role emphasizes production-grade coding, with strong skills in Go, C/C++, Rust, or Scala, and familiarity with time-series databases and streaming

Qualifications

  • BS (or equivalent) plus 5+ years experience, MS 3+ years, or PhD with 1+ year in CS/Stats.
  • Strong mathematical foundation: statistics, probability, linear algebra, and algorithm analysis.
  • Proven experience implementing and optimizing ML algorithms in production — coding-first role; strong implementation skills required.
  • Strong programming skills in Go, C/C++, Rust, or Scala; Python working knowledge is a plus.
  • Familiarity with time-series databases and streaming data architectures
  • Ability to work independently and navigate ambiguity in a fast-paced engineering environment

Responsibilities

  • Implement production ML algorithms in Go — optimized for real-time streaming pipelines operating at massive scale under strict resource constraints
  • Design and develop new ML algorithms where needed: anomaly detection, health scoring, and predictive analytics on high-volume time-series telemetry from GPU and network infrastructure
  • Improve and extend existing algorithms and experiment with new approaches suited to real-time streaming constraints
  • Build and maintain end-to-end ML pipelines — from data ingestion and schema design through model inference — optimized for on-premises, latency-sensitive deployments
  • Partner with the Data Science team on algorithm design, prototype evaluation, and translating research findings into platform requirements
  • Collaborate with teams to integrate ML into GPU networking infrastructure for AI workloads

Skills

Go
C/C++
Rust
Python

Education

BS in Computer Science or related field

Tools

Kafka

Job description

## Machine Learning Systems Engineer, NetworkingApplylocations: US, CA, Santa Claratime type: Full timeposted on: Posted Todayjob requisition id: JR2018261Join our team of innovative engineers who are building an AI Data Center AIOps platform that turns raw, high-volume telemetry into reliable, job-centric insights and automation for GPU fleets. As an ML Engineer on this team, you'll design and implement ML algorithms that run in real-time streaming pipelines, detecting anomalies and surfacing insights across massive-scale infrastructure before they impact AI training and inference.The core challenge of this role is building ML algorithms that are simultaneously accurate and efficient —processing millions of telemetry streams in real time within tight CPU and memory budgets. You'll need both the data science depth to design and validate algorithms and the engineering discipline to implement them in production at scale.**What you'll be doing:*** Implement production ML algorithms in Go — optimized for real-time streaming pipelines operating at massive scale under strict resource constraints* Design and develop new ML algorithms where needed: anomaly detection, health scoring, and predictive analytics on high-volume time-series telemetry from GPU and network infrastructure* Improve and extend existing algorithms and experiment with new approaches suited to real-time streaming constraints* Build and maintain end-to-end ML pipelines — from data ingestion and schema design through model inference — optimized for on-premises, latency-sensitive deployments* Partner with the Data Science team on algorithm design, prototype evaluation, and translating research findings into platform requirements**What we need to see:*** A BS (or equivalent experience) and 5+ years of experience, MS and 3+ years, or PhD with 1+ years in Computer Science, Statistics, or a related field* Strong mathematical foundation: statistics, probability, linear algebra, and algorithm analysis* Proven experience implementing and optimizing ML algorithms in production — this is a coding-first role; strong implementation skills are required* Strong programming skills in one or more of Go, C/C++, Rust, or Scala; Python working knowledge is a plus* Familiarity with time-series databases and streaming data architectures* Ability to work independently and navigate ambiguity in a fast-paced engineering environment**Ways to stand out from the crowd:*** Data Science background with hands-on experience building and validating ML models — bridging research and production implementation* Experience implementing ML algorithms directly in systems languages for latency-sensitive or resource-constrained environments* Research experience: knowing the latest ML literature and translating advances into practical improvements* Experience with Kafka-based streaming pipelines and real-time feature engineering at scaleWith competitive salaries and a generous benefits package, we are widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and hardworking people in the world working for us and, due to unprecedented growth, our exclusive engineering teams are rapidly growing. If you're a creative and autonomous engineer with a real passion for technology, we want to hear from you.Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 152,000 USD - 241,500 USD for Level 3, and 184,000 USD - 287,500 USD for Level 4.You will also be eligible for equity and benefits.Applications for this job will be accepted at least until August 15, 2026.This posting is for an existing vacancy.NVIDIA uses AI tools in its recruiting processes.NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.NVIDIA pioneered accelerated computing. Today, our AI infrastructure powers global intelligence, transforming every industry.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Machine Learning Systems Engineer, Networking
Machine Learning Systems Engineer, Networking

NVIDIA AI • Santa Clara (CA)

On-site
USD 152,000 - 288,000
Equity
Benefits
Senior Software Engineer, At Scale Compute Analysis
Senior Software Engineer, At Scale Compute Analysis

NVIDIA Corporation • Santa Clara (CA)

Hybrid
USD 152,000 - 242,000
Health coverage
Dental + Vision
401(k) match
+6
Senior AI Workflow Engineer
Senior AI Workflow Engineer

NVIDIA Corporation • Santa Clara (CA), Northern (KY)

Hybrid
USD 184,000 - 357,000
Equity compensation
Health insurance
Relocation support
Senior Platform AI Engineer
Senior Platform AI Engineer

NVIDIA Corporation • Santa Clara (CA)

Hybrid
USD 184,000 - 357,000
Equity
Comprehensive benefits package
Senior Solutions Architect, AI Hyperscalers
Senior Solutions Architect, AI Hyperscalers

NVIDIA Corporation • Santa Clara (CA)

Hybrid
USD 184,000 - 288,000
Equity
Comprehensive benefits
Machine Learning Systems Engineer, Networking
Machine Learning Systems Engineer, Networking

NVIDIA • Santa Clara (CA)

On-site
USD 152,000 - 287,500
Senior Software Engineer, Machine Learning Inference
Senior Software Engineer, Machine Learning Inference

NVIDIA Corporation • Santa Clara (CA)

Hybrid
USD 152,000 - 288,000
Equity
Benefits
Senior Platform Telemetry Engineer
Senior Platform Telemetry Engineer

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 152,000 - 288,000
Senior AI Compute Engineer - NVIS
Senior AI Compute Engineer - NVIS

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 148,000 - 236,000
Senior Software Architect, AI Systems and Networking
Senior Software Architect, AI Systems and Networking

NVIDIA Corporation • Austin (TX)

On-site
USD 184,000 - 357,000
Generous benefits package
Competitive salaries