Engineering Manager, Agentic GenAI Platform

NVIDIA Corporation

Santa Clara (CA)

On-site

USD 224,000 - 431,250

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

NVIDIA is seeking an Engineering Manager to lead a team building an agentic platform for observing, debugging, and optimizing GenAI models deployed at scale, with visibility into model behavior and performance across large-scale workloads.

You will guide development across logs, traces, metrics, and production AI systems, drive reliability, latency improvements, and cost optimization, and collaborate with stakeholders to align priorities.

Qualifications

  • BS/MS/PhD in CS or CE or equivalent experience.
  • 8+ years of software engineering experience with 3+ years in management or technical leadership.
  • Experience leading teams building large-scale distributed systems and observability platforms.
  • Strong knowledge of LLM/VLM inference, production performance challenges.
  • Experience with logs, metrics, traces, profiling, alerting, dashboards, or debugging workflows.
  • Strong programming, debugging, performance analysis, and test design skills.
  • Ability to align technical work with product and business goals.
  • Excellent communication and collaboration skills.

Responsibilities

  • Lead, mentor, and grow a team building an agentic platform for monitoring and improving large-scale LLMs and VLMs in production.
  • Build systems that collect, correlate, and analyze telemetry across inference servers, GPUs, schedulers, model runtimes, and customer-facing APIs.
  • Develop agentic workflows to identify root causes, explain regressions, and suggest performance optimizations.
  • Collaborate with internal customers and business units to align priorities and deliver production-grade capabilities.

Skills

Team leadership
Distributed systems
Observability platforms
LLM/VLM knowledge
Programming and debugging
Communication skills

Education

BSc, MS, or PhD in CS/CE

Tools

OpenTelemetry
Prometheus
Grafana
Jaeger
ClickHouse
Elastic

Job description

NVIDIA is seeking an Engineering Manager to lead the development of an agentic platform for observing, debugging, and optimizing GenAI models deployed at scale. In this role, you will lead a team building an agentic platform that provides visibility into model behavior, inference performance, reliability, and cost across large-scale GenAI workloads. The platform will capture, correlate, and analyze logs, traces, metrics, and performance signals across large‑scale LLM and VLM deployments. It will help engineers understand model‑serving behavior, find regressions, optimize latency and throughput, and improve the reliability of GenAI systems in production.

What You’ll Be Doing
  • Lead, mentor, and grow a team building an agentic platform for monitoring and improving large‑scale LLMs and VLMs in production.
  • Build systems that collect, correlate, and analyze telemetry across inference servers, GPUs, schedulers, model runtimes, and customer‑facing APIs.
  • Develop agentic workflows that help engineers identify root causes, explain regressions, and recommend performance optimizations.
  • Collaborate with internal customers and business units to align priorities and deliver production‑grade platform capabilities.
What We Need To See
  • BSc, MS, or PhD in Computer Science, Computer Engineering, or equivalent experience.
  • 8+ years of relevant software engineering experience, including 3+ years in engineering management or technical leadership.
  • Experience leading software engineering teams building large‑scale distributed systems, observability platforms, ML infrastructure, or production AI systems.
  • Strong understanding of LLM/VLM inference systems, deployment patterns, and production performance challenges.
  • Experience with logs, metrics, traces, profiling, alerting, dashboards, or incident/debugging workflows.
  • Strong programming, debugging, performance analysis, and test design skills.
  • Ability to work across organizations and align technical priorities with product and business goals.
  • Excellent communication and collaboration skills.
Ways To Stand Out From The Crowd
  • Background in GPU performance analysis, distributed inference, model serving optimization, or reliability engineering.
  • Experience building observability or telemetry platforms for AI, ML, cloud, or distributed infrastructure.
  • Experience with OpenTelemetry, Prometheus, Grafana, Jaeger, ClickHouse, Elastic, or similar observability tools.
  • Experience building agentic systems that reason over logs, traces, performance data, incidents, or operational workflows.
  • Hands‑on experience with production GenAI serving systems and metrics such as TTFT, TPOT, throughput, queueing delay, GPU utilization, KV cache pressure, error rates, and cost per token.
Compensation and Benefits

Base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is $224,000 – $356,500 USD for Level3, and $272,000 – $431,250 USD for Level4. You will also be eligible for equity and benefits.

#LI-Hybrid

Equal Employment Opportunity

NVIDIA is committed to fostering an inclusive work environment and is proud to be an equal opportunity employer. We do not discriminate based on race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status, or any other characteristic protected by law.

Applications for this job will be accepted at least until July20,2026.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Engineering Manager, Agentic GenAI Platform
Engineering Manager, Agentic GenAI Platform

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 224,000 - 432,000
Equity
Benefits
Engineering Manager, Agentic GenAI Platform
Engineering Manager, Agentic GenAI Platform

NVIDIA • Santa Clara (CA)

On-site
USD 224,000 - 432,000
Equity
Benefits
Senior Developer Technology Engineer - Edge Agentic AI
Senior Developer Technology Engineer - Edge Agentic AI

Segment (Twilio) • Santa Clara (CA)

On-site
USD 152,000 - 288,000
Equity
Benefits package
Senior Solutions Architect, Gen AI
Senior Solutions Architect, Gen AI

NVIDIA Gruppe • California (MO)

On-site
USD 184,000 - 288,000
Equity
Benefits
Senior Software Engineer – Platform Engineering
Senior Software Engineer – Platform Engineering

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 200,000 - 322,000
Equity
Benefits
Senior Software Engineer, Agentic AI
Senior Software Engineer, Agentic AI

NVIDIA • Santa Clara (CA)

On-site
USD 152,000 - 242,000
Equity
Benefits
Senior Software Engineer, Agentic AI
Senior Software Engineer, Agentic AI

NVIDIA • Redmond (WA)

On-site
USD 184,000 - 288,000
Equity
Benefits
Engineering Manager, Deep Learning Inference
Engineering Manager, Deep Learning Inference

Socket.dev • Santa Clara (UT)

On-site
USD 224,000 - 357,000
Equity
Benefits package
Engagement Tech Lead, Agentic AI
Engagement Tech Lead, Agentic AI

Nvidia Corporation • Santa Clara (CA)

On-site
USD 224,000 - 431,250
Equity
Benefits
Senior System Software Engineer, Agentic Inference - Dynamo
Senior System Software Engineer, Agentic Inference - Dynamo

NVIDIA • Santa Clara (CA)

Hybrid
USD 272,000 - 431,250
Equity
Benefits
Hybrid work