Realtime ML Systems Engineer, Networking & AIOps

NVIDIA Corporation

Santa Clara, Northern (CA, KY)

Hybrid

USD 152,000 - 288,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

NVIDIA Corporation is seeking a Machine Learning Systems Engineer, Networking to design and implement ML algorithms for real-time telemetry pipelines across GPU fleets. You will build end-to-end ML pipelines and optimize performance within strict CPU/memory budgets, enabling proactive AI training and inference across massive-scale infrastructure.

The role emphasizes production-grade coding, with strong skills in Go, C/C++, Rust, or Scala, and familiarity with time-series databases and streaming

Qualifications

  • BS (or equivalent) plus 5+ years experience, MS 3+ years, or PhD with 1+ year in CS/Stats.
  • Strong mathematical foundation: statistics, probability, linear algebra, and algorithm analysis.
  • Proven experience implementing and optimizing ML algorithms in production — coding-first role; strong implementation skills required.
  • Strong programming skills in Go, C/C++, Rust, or Scala; Python working knowledge is a plus.
  • Familiarity with time-series databases and streaming data architectures
  • Ability to work independently and navigate ambiguity in a fast-paced engineering environment

Responsibilities

  • Implement production ML algorithms in Go — optimized for real-time streaming pipelines operating at massive scale under strict resource constraints
  • Design and develop new ML algorithms where needed: anomaly detection, health scoring, and predictive analytics on high-volume time-series telemetry from GPU and network infrastructure
  • Improve and extend existing algorithms and experiment with new approaches suited to real-time streaming constraints
  • Build and maintain end-to-end ML pipelines — from data ingestion and schema design through model inference — optimized for on-premises, latency-sensitive deployments
  • Partner with the Data Science team on algorithm design, prototype evaluation, and translating research findings into platform requirements
  • Collaborate with teams to integrate ML into GPU networking infrastructure for AI workloads

Skills

Go
C/C++
Rust
Python

Education

BS in Computer Science or related field

Tools

Kafka

Job description

NVIDIA Corporation is seeking a Machine Learning Systems Engineer, Networking to design and implement ML algorithms for real-time telemetry pipelines across GPU fleets. You will build end-to-end ML pipelines and optimize performance within strict CPU/memory budgets, enabling proactive AI training and inference across massive-scale infrastructure.

The role emphasizes production-grade coding, with strong skills in Go, C/C++, Rust, or Scala, and familiarity with time-series databases and streaming

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Real-Time ML Systems Engineer — Anomaly Detection & Streaming
Real-Time ML Systems Engineer — Anomaly Detection & Streaming

NVIDIA • Santa Clara (CA)

On-site
USD 152,000 - 287,500
Go-powered ML Systems Engineer - Real-Time AI for GPUs
Go-powered ML Systems Engineer - Real-Time AI for GPUs

NVIDIA AI • Santa Clara (CA)

On-site
USD 152,000 - 288,000
Equity
Benefits
Machine Learning Systems Engineer, Networking
Machine Learning Systems Engineer, Networking

NVIDIA • Santa Clara (CA)

On-site
USD 152,000 - 287,500
Machine Learning Systems Engineer, Networking
Machine Learning Systems Engineer, Networking

NVIDIA AI • Santa Clara (CA)

On-site
USD 152,000 - 288,000
Equity
Benefits
Machine Learning Systems Engineer, Networking
Machine Learning Systems Engineer, Networking

NVIDIA Corporation • Santa Clara (CA), Northern (KY)

Hybrid
USD 152,000 - 288,000
Senior ML Engineer - Real-Time Networking AI Leader
Senior ML Engineer - Real-Time Networking AI Leader

Extreme Networks • Seattle (WA)

On-site
USD 180,000 - 240,000
Senior ML Graphics Engineer - Real-Time AI for Gaming
Senior ML Graphics Engineer - Real-Time AI for Gaming

Nvidia Corporation • Santa Clara (CA)

On-site
USD 272,000 - 432,000
Equity
Benefits
Senior ML Engineer — Real-Time AI for Networking
Senior ML Engineer — Real-Time AI for Networking

Extremenetworks • Seattle (WA)

On-site
USD 180,000 - 240,000
Senior Principal ML Engineer: Real-Time AI for Networks
Senior Principal ML Engineer: Real-Time AI for Networks

Socket.dev • Seattle (WA)

On-site
USD 180,000 - 250,000
Lead ML Systems Engineer — Distributed GPU Training & Infra
Lead ML Systems Engineer — Distributed GPU Training & Infra

Nvidia Corporation • Santa Clara (CA)

On-site
USD 224,000 - 431,000
Equity
Benefits