Machine Learning Systems Engineer, Networking

NVIDIA

Santa Clara (CA)

On-site

USD 152,000 - 287,500

Full time

10 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

NVIDIA seeks an ML Engineer to design and implement real-time ML algorithms that process millions of telemetry streams for an AI Data Center AIOps platform. You will build production-grade models and pipelines that detect anomalies and surface insights under tight CPU and memory budgets.

You’ll code in Go, C/C++, Rust or Scala, collaborate with Data Science to translate research into platform requirements, and contribute to scalable on‑premises deployments.

Qualifications

  • BS (or equivalent) in Computer Science, Statistics, or related field with 5+ years experience; MS with 3+ years; or PhD with 1+ year.
  • Strong foundation in statistics, probability, linear algebra and algorithms.
  • Production ML experience; strong coding skills in one of Go, C/C++, Rust, or Scala.

Responsibilities

  • Implement production ML algorithms in Go for real-time streaming pipelines.
  • Design and develop anomaly detection, health scoring, and predictive analytics on high-volume telemetry.
  • Improve and extend existing ML algorithms for latency-sensitive deployments.
  • Build end-to-end ML pipelines from data ingestion to model inference for on-premises systems.
  • Collaborate with Data Science to translate research into platform requirements.

Skills

Go
C/C++
Rust
Scala
Python

Education

Bachelor's degree in CS/Math/related field
MS in CS/Statistics
PhD in CS/Statistics

Tools

Kafka
Time-series databases
Streaming frameworks

Job description

Join our team of innovative engineers that are building an AI Data Center AIOps platform that turns raw, high-volume telemetry into reliable, job-centric insights and automation for GPU fleets. As an ML Engineer on this team, you'll design and implement ML algorithms that run in real-time streaming pipelines, detecting anomalies and surfacing insights across massive-scale infrastructure before they impact AI training and inference.

The core challenge of this role is building ML algorithms that are simultaneously accurate and efficient —processing millions of telemetry streams in real time within tight CPU and memory budgets. You'll need both the data science depth to design and validate algorithms and the engineering discipline to implement them in production at scale.

What you’ll be doing:
  • Implement production ML algorithms in Go — optimized for real-time streaming pipelines operating at massive scale under strict resource constraints
  • Design and develop new ML algorithms where needed: anomaly detection, health scoring, and predictive analytics on high-volume time‑series telemetry from GPU and network infrastructure
  • Improve and extend existing algorithms and experiment with new approaches suited to real‑time streaming constraints
  • Build and maintain end‑to‑end ML pipelines — from data ingestion and schema design through model inference — optimized for on‑premises, latency‑sensitive deployments
  • Partner with the Data Science team on algorithm design, prototype evaluation, and translating research findings into platform requirements
What we need to see:
  • A BS (or equivalent experience) and 5+ years of experience, MS and 3+ years, or PhD with 1+ years in Computer Science, Statistics, or a related field
  • Strong mathematical foundation: statistics, probability, linear algebra, and algorithm analysis
  • Proven experience implementing and optimizing ML algorithms in production — this is a coding‑first role; strong implementation skills are required
  • Strong programming skills in one or more of Go, C/C++, Rust, or Scala; Python working knowledge is a plus
  • Familiarity with time‑series databases and streaming data architectures
  • Ability to work independently and navigate ambiguity in a fast‑paced engineering environment
Ways to stand out from the crowd:
  • Data Science background with hands‑on experience building and validating ML models — bridging research and production implementation
  • Experience implementing ML algorithms directly in systems languages for latency‑sensitive or resource‑constrained environments
  • Research experience: knowing the latest ML literature and translating advances into practical improvements
  • Experience with Kafka‑based streaming pipelines and real‑time feature engineering at scale

With competitive salaries and a generous benefits package, we are widely considered to be one of the technology world’s most desirable employers. We have some of the most forward‑thinking and hardworking people in the world working for us and, due to unprecedented growth, our exclusive engineering teams are rapidly growing. If you're a creative and autonomous engineer with a real passion for technology, we want to hear from you.

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 152,000 USD – 241,500 USD for Level 3, and 184,000 USD – 287,500 USD for Level 4.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until July 14, 2026.

NVIDIA uses AI tools in its recruiting processes.

NVIDIA is committed to fostering an inclusive work environment and proud to be an equal‑opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Machine Learning Systems Engineer, Networking
Machine Learning Systems Engineer, Networking

NVIDIA AI • Santa Clara (CA)

On-site
USD 152,000 - 288,000
Equity
Benefits
Machine Learning Systems Engineer, Networking
Machine Learning Systems Engineer, Networking

NVIDIA Corporation • Santa Clara (CA), Northern (KY)

Hybrid
USD 152,000 - 288,000
Senior AI and HPC Observability Engineer
Senior AI and HPC Observability Engineer

NVIDIA • Santa Clara (CA)

On-site
USD 152,000 - 242,000
Equity
Benefits
Applied AI Engineer
Applied AI Engineer

2100 NVIDIA USA • California (MO)

On-site
USD 152,000 - 288,000
Equity
Benefits
Senior Software Engineer, AI Inference Systems
Senior Software Engineer, AI Inference Systems

NVIDIA • Santa Clara (CA)

Hybrid
USD 184,000 - 357,000
Equity
Benefits
Applied AI Engineer
Applied AI Engineer

NVIDIA AI • Santa Clara (CA)

On-site
USD 152,000 - 288,000
Equity
Benefits package
Senior Site Reliability Engineer, AIOPs
Senior Site Reliability Engineer, AIOPs

NVIDIA • Santa Clara (CA)

On-site
USD 148,000 - 276,000
Equity
Benefits
Senior Deep Learning Algorithm Engineer
Senior Deep Learning Algorithm Engineer

NVIDIA AI • Santa Clara (CA)

On-site
USD 152,000 - 288,000
Equity
Benefits
Senior AI Performance and Efficiency Engineer
Senior AI Performance and Efficiency Engineer

Nvidia Corporation • Santa Clara (CA)

On-site
USD 184,000 - 288,000
Competitive salaries
Comprehensive benefits package
Equity eligibility
Applied AI Engineer
Applied AI Engineer

NVIDIA • Santa Clara (CA)

On-site
USD 152,000 - 288,000
Equity
Benefits package
Hybrid work model