Distinguished Engineer, Scaled Out Inferencing

Nvidia Corporation

Santa Clara (CA)

On-site

USD 320,000 - 489,000

Full time

12 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
Benefits package

Job summary

NVIDIA is seeking a senior AI infrastructure leader with extensive experience in large-scale inference orchestration and secure, highly available distributed systems. You will steer cross-domain optimization across hardware and software, guiding multi-tenant deployments and cloud/datacenter strategies.

Lead architecture and lifecycle management while engaging with customers, partners, and executive leadership to ensure NVIDIA’s solutions set industry standards for performance and reliability.

Qualifications

  • 16+ years in technical roles with focus on AI infrastructure and large-scale inference orchestration
  • 7-10+ years of leadership experience
  • BS/MS or higher in systems / software engineering or related fields
  • Deep technical expertise in GPU architecture, hardware acceleration, and low-level performance tuning (CUDA, kernels)
  • Proven success delivering high-impact, transparent resource/utilization and performance insights
  • Technical leadership and cross-functional alignment at senior levels
  • Strong communication and teamwork to engage engineering, customers, and partners

Responsibilities

  • Architect distributed pipelines for high-throughput, low-latency inference systems
  • Drive hardware-software co-optimization and performance tuning
  • Guide open source projects and ecosystem initiatives (Dynamo, TensorRT-LLM, vLLM, SGLang, Ray)
  • Lead model lifecycle management including automated deployment and scaling
  • Collaborate with customers and partners to define industry-leading standards
  • Oversee the full software and system lifecycle from ideation to evolution

Skills

GPU architecture
Hardware acceleration
Low-level tuning
Cloud-native architecture
Leadership
Cross-functional collaboration
Strategic thinking

Education

BS/MS or higher in systems/ software engineering

Tools

CUDA
Kernels
Linux
Kubernetes
Ray

Job description

NVIDIA is leading the industry in delivering accelerated computing in cloud and enterprise environments. We’re a team of innovative engineers dedicated to solving some of the world’s biggest challenges, constantly driving advancements, and impacting millions of lives worldwide!

What You’ll Be Doing:
  • Various Architectural Work: Architect distributed pipelines, define and drive the technical implementation of high-throughput, low-latency, distributed inference systems to support massive-scale AI workloads.
  • Collaborate on Cross Domain Disciplines: Hardware-software co-optimization, drive performance tuning at the kernel and driver level, optimizing GPU resource management and hardware acceleration for production-grade model serving.
  • Collaborate on Open Source and Ecosystem Projects: Guide and influence open source projects Dynamo, TensorRT-LLM, and ecosystem projects (vLLM, SGLang, Linux, Kubernetes, Ray) to bring state of the art inferencing on NVIDIA accelerated hardware.
  • Accelerate Integration: Orchestrate model lifecycles, lead the strategy for full-lifecycle model management, including automated deployment, versioning, and intelligent scaling across varied cloud and datacenter environments.
  • Engage Stakeholders: Collaborate with customers, infrastructure providers, and partners to ensure NVIDIA’s solutions set the industry standard for performance and availability.
  • Full Software and System Lifecycle: From ideation to architecture, design, development, deployment, operations, and full lifecycle management, lead all technical aspects of planning and continuous evolution of a large technical scope.
What We Need to See:
  • 16+ overall years in technical roles with a recent long-term focus on AI infrastructure and more recent direct experience in large-scale inference orchestration. Proven track record building secure, highly available, and durable production distributed systems.
  • 7-10+ years of leadership experience
  • BS/MS or higher or equivalent experience in systems / software engineering, or related engineering fields
  • Deep Technical Expertise: Proficiency in GPU architecture, hardware acceleration, and low-level performance tuning (CUDA, kernels) alongside cloud-native architectures for multi-tenant model serving.
  • Proven success delivering high-impact technically complex solutions that achieve high levels of transparency into resource utilization, performance, and operational insights.
  • Technical Leadership: Develop and advance consensus and organizational alignment across technical leadership and the highest level of senior corporate leadership. Ability to synthesize cross-functional needs into architecture and design while guiding internal execution across diverse teams.
  • Communication and Teamwork: Strong collaboration and influence skills, capable of leading engineering engagement, communicating with peers, partners, and working with high performance and accelerated computing customers.
Ways to Stand Out from the Crowd:
  • Application of Artificial Intelligence: Real world experience building the systems to support AI/ML workloads.
  • Industry Expertise: Direct experience in designing, developing, delivering and operating secure, highly available, scaled out systems in enterprise and cloud environments.
  • Engineering Enablement: Demonstrated history of creating scalable processes and extensible systems that facilitate cross-functional collaboration and operations at scale.
  • Open Source Collaboration: Familiarity with open source ecosystems and projects (e.g. Dynamo, TensorRT-LLM, vLLM, SGLang, Ray). Ability to collaborate and influence in open source project governance to represent NVIDIA, customers, and partners interests in technical alignment and direction.

NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High-Performance Computing and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. We have some of the most forward-thinking and hardworking people on the planet working for us. If you're creative, passionate and self-motivated, we want to hear from you!

With competitive salaries and a generous benefits package (www.nvidiabenefits.com), we are widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and hardworking people in the world working for us and, due to outstanding growth, our best-in-class engineering teams are rapidly growing. If you're a creative and autonomous engineer with a real passion for technology, we want to hear from you!

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 320,000 USD - 488,750 USD.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until August 15, 2026.

This posting is for an existing vacancy.

NVIDIA uses AI tools in its recruiting processes.

NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Distinguished Engineer, Scaled Out Inferencing
Distinguished Engineer, Scaled Out Inferencing

NVIDIA • California (MO)

On-site
USD 320,000 - 489,000
Equity
Benefits
Distinguished Engineer, Scaled Out Inferencing
Distinguished Engineer, Scaled Out Inferencing

NVIDIA • Santa Clara (CA)

On-site
USD 320,000 - 489,000
Distinguished Engineer, Scaled Out Inferencing
Distinguished Engineer, Scaled Out Inferencing

2100 NVIDIA USA • Santa Clara (CA)

On-site
USD 320,000 - 489,000
Distinguished Engineer, Scaled Out Inferencing
Distinguished Engineer, Scaled Out Inferencing

Nvidia Corporation in • Santa Clara (CA)

On-site
USD 320,000 - 489,000
Equity
Benefits package
Distinguished Engineer, System Software Integration
Distinguished Engineer, System Software Integration

Nvidia Corporation • Santa Clara (CA)

On-site
USD 320,000 - 489,000
Equity
Benefits package
Distinguished Engineer, System Software Integration
Distinguished Engineer, System Software Integration

NVIDIA • United States

On-site
USD 320,000 - 489,000
Distinguished Engineer, System Software Integration
Distinguished Engineer, System Software Integration

Nvidia Corporation in • Santa Clara (CA)

On-site
USD 320,000 - 489,000
Equity compensation
Benefits package
Competitive salary
Distinguished Engineer, System Software Integration
Distinguished Engineer, System Software Integration

2100 NVIDIA USA • California (MO)

On-site
USD 320,000 - 489,000
Distinguished Engineer, System Software Integration
Distinguished Engineer, System Software Integration

NVIDIA Gruppe • California (MO)

On-site
USD 320,000 - 489,000
Equity
Benefits package
Distinguished Engineer, System Software Integration
Distinguished Engineer, System Software Integration

NVIDIA AI • California (MO)

On-site
USD 320,000 - 489,000
Equity
Benefits