Engineering Manager, Agentic GenAI Platform

NVIDIA Gruppe

Santa Clara (CA)

On-site

USD 224,000 - 431,250

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
Benefits

Job summary

NVIDIA is seeking an Engineering Manager to lead the development of an agentic platform for observing, debugging, and optimizing GenAI models deployed at scale. You will guide a team to build visibility into model behavior, inference performance, reliability, and cost across large-scale workloads.

This role spans across the NVIDIA AI software stack, collaborating with teams in inference serving, model optimization, distributed systems, and production operations.

Qualifications

  • BSc, MS, or PhD in Computer Science, Computer Engineering, or equivalent experience.
  • 8+ years of relevant software engineering experience, including 3+ years in management.
  • Experience leading software teams building large-scale distributed systems and observability platforms.
  • Strong understanding of LLM/VLM inference and production performance challenges.
  • Experience with logs, metrics, traces, profiling, alerting, dashboards, or incident workflows.
  • Excellent communication and collaboration skills.

Responsibilities

  • Lead, mentor, and grow a team building an agentic platform for monitoring and improving large-scale LLMs and VLMs in production.
  • Build systems that collect, correlate, and analyze telemetry across inference servers, GPUs, schedulers, model runtimes, and customer-facing APIs.
  • Develop agentic workflows that help engineers identify root causes, explain regressions, and recommend performance optimizations.
  • Collaborate with internal customers and business units to align priorities and deliver production grade platform capabilities.

Skills

Team leadership
Distributed systems
Observability platforms
ML infrastructure
Performance analysis
Communication
Cross-functional collaboration

Education

BSc/MS/PhD in CS/CE

Tools

OpenTelemetry
Prometheus
Grafana
Jaeger
ClickHouse
Elastic

Job description

NVIDIA is seeking an Engineering Manager to lead the development of an agentic platform for observing, debugging, and optimizing GenAI models deployed at scale. In this role, you will lead a team building an agentic platform that provides visibility into model behavior, inference performance, reliability, and cost across large-scale GenAI workloads. This agentic platform will capture, correlate, and analyze logs, traces, metrics, and performance signals across large scale LLM and VLM deployments. It will help engineers understand model-serving behavior, find regressions, optimize latency and throughput, and improve the reliability of GenAI systems in production.

You will work across the NVIDIA AI software stack with teams focused on inference serving, model optimization, distributed systems, GPU performance, and production operations. This is a highly cross‑functional role for someone who understands deep learning systems, observability, and large‑scale software platforms, and who is excited about building agentic workflows that help teams reason over complex telemetry and performance data.

What You’ll Be Doing
  • Lead, mentor, and grow a team building an agentic platform for monitoring and improving large‑scale LLMs and VLMs in production.
  • Build systems that collect, correlate, and analyze telemetry across inference servers, GPUs, schedulers, model runtimes, and customer‑facing APIs.
  • Develop agentic workflows that help engineers identify root causes, explain regressions, and recommend performance optimizations.
  • Collaborate with internal customers and business units to align priorities and deliver production grade platform capabilities.
What We Need To See
  • BSc, MS, or PhD in Computer Science, Computer Engineering, or equivalent experience.
  • 8+ years of relevant software engineering experience, including 3+ years in engineering management or technical leadership.
  • Experience leading software engineering teams building large‑scale distributed systems, observability platforms, ML infrastructure, or production AI systems.
  • Strong understanding of LLM/VLM inference systems, deployment patterns, and production performance challenges.
  • Experience with logs, metrics, traces, profiling, alerting, dashboards, or incident/debugging workflows.
  • Strong programming, debugging, performance analysis, and test design skills.
  • Ability to work across organizations and align technical priorities with product and business goals.
  • Excellent communication and collaboration skills.
Ways To Stand Out From The Crowd
  • Background in GPU performance analysis, distributed inference, model serving optimization, or reliability engineering.
  • Experience building observability or telemetry platforms for AI, ML, cloud, or distributed infrastructure.
  • Experience with OpenTelemetry, Prometheus, Grafana, Jaeger, ClickHouse, Elastic, or similar observability tools.
  • Experience building agentic systems that reason over logs, traces, performance data, incidents, or operational workflows.
  • Hands‑on experience with production GenAI serving systems and metrics such as TTFT, TPOT, throughput, queueing delay, GPU utilization, KV cache pressure, error rates, and cost per token.

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is $224,000 USD – $356,500 USD for Level 3, and $272,000 USD – $431,250 USD for Level 4.

You will also be eligible for equity and benefits.

NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Engineering Manager, Agentic GenAI Platform
Engineering Manager, Agentic GenAI Platform

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 224,000 - 431,250
Engineering Manager, Agentic GenAI Platform
Engineering Manager, Agentic GenAI Platform

NVIDIA • Santa Clara (CA)

On-site
USD 224,000 - 432,000
Equity
Benefits
Senior Developer Technology Engineer - Edge Agentic AI
Senior Developer Technology Engineer - Edge Agentic AI

Segment (Twilio) • Santa Clara (CA)

On-site
USD 152,000 - 288,000
Equity
Benefits package
Senior Solutions Architect, Gen AI
Senior Solutions Architect, Gen AI

NVIDIA Gruppe • California (MO)

On-site
USD 184,000 - 288,000
Equity
Benefits
Senior Software Engineer – Platform Engineering
Senior Software Engineer – Platform Engineering

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 200,000 - 322,000
Equity
Benefits
Senior Software Engineer, Agentic AI
Senior Software Engineer, Agentic AI

NVIDIA • Santa Clara (CA)

On-site
USD 152,000 - 242,000
Equity
Benefits
Senior Software Engineer, Agentic AI
Senior Software Engineer, Agentic AI

NVIDIA • Redmond (WA)

On-site
USD 184,000 - 288,000
Equity
Benefits
Engineering Manager, Deep Learning Inference
Engineering Manager, Deep Learning Inference

Socket.dev • Santa Clara (UT)

On-site
USD 224,000 - 357,000
Equity
Benefits package
ML and Agentic Systems Engineer
ML and Agentic Systems Engineer

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 224,000 - 357,000
Equity
Benefits
Engineering Manager, Deep Learning Inference
Engineering Manager, Deep Learning Inference

NVIDIA • Washington

On-site
USD 184,000 - 357,000
Equity
Benefits package