Senior Solutions Architect, GPU Cloud GenAI – Infrastructure

NVIDIA

Mumbai City

On-site

INR 3,000,000 - 4,200,000

Full time

5 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

NVIDIA is seeking an experienced Senior Solutions Architect & Engineer to lead GPU cloud infrastructure for GenAI workloads in Mumbai. You will design scalable IaaS/PaaS/SaaS layers across large GPU clusters, build multi-tenant platforms with robust APIs, and implement Kubernetes and Slurm pipelines for high-performance AI workloads.

Engage with executives and customer teams to drive platform adoption and cost-efficient operation.

Qualifications

  • 5+ years designing and operating large GPU clusters (100+ nodes).
  • Proficient in IaaS, PaaS, SaaS platform layers and infrastructure foundations.
  • Strong Kubernetes, Slurm, and multi-tenant orchestration experience.
  • Experience with IaC, CI/CD, observability stacks, and scripting.

Responsibilities

  • Design scalable IaaS, PaaS, and SaaS layers for large GPU clusters.
  • Build multi-tenant GPU cloud platforms with production-grade APIs.
  • Develop pipelines using Kubernetes and Slurm for HPC/AI workloads.
  • Implement best practices for scheduling, isolation, and observability.
  • Advise on deploying generative AI workloads and enterprise deployments.
  • Collaborate with teams to scale GPU clusters across on‑prem and hybrid environments.

Skills

GPU clusters
Kubernetes
Slurm
Terraform
Helm
Python
Go/C++

Education

Bachelor's degree in CS/CE or equivalent

Tools

Terraform
Helm
Ansible
Prometheus
Grafana
DC/GC

Job description

Senior Solutions Architect, GPU Cloud GenAI - Infrastructure

NVIDIA is seeking an experienced Solutions Architect & Engineer (SAE) with deep expertise in large-scale GPU cluster infrastructure and generative AI enablement. As a pivotal member of our Infrastructure and Platform Engineering team, you will architect and build the GPU cloud platforms (IaaS, PaaS, SaaS) that power the world's most demanding AI workloads. This position sits at the intersection of large-scale infrastructure engineering and applied AI, requiring both technical depth in platform development and the ability to guide enterprise customers through complex GPU infrastructure deployments.

The work location for this role is in Mumbai.

What you will be doing:

Design and architect scalable IaaS, PaaS, and SaaS layers for large-scale GPU cluster environments (32+ HGX/DGX nodes), spanning compute, networking, and storage orchestration.

Build multi-tenant GPU cloud platforms with production-grade APIs, control planes, and platform services that abstract infrastructure complexity for end users and application teams.

Develop cluster orchestration pipelines using Kubernetes (GPU operators, device plugins, multi-tenancy) and Slurm, optimizing for performance, reliability, and resource efficiency at scale.

Define and implement best practices for GPU resource scheduling, isolation, quota management, and observability, ensuring secure multi-tenant isolation and compliance.

Advise customers on deploying and scaling generative AI workloads (LLMs, MLLMs, RAG pipelines) on your infrastructure platforms, translating AI requirements into infrastructure specifications.

Engage with C-level executives and infrastructure teams to understand requirements, deploy GPU clusters across on-premises and hybrid cloud environments, and drive platform adoption.

Collaborate with NVIDIA engineering teams to resolve deep infrastructure bugs, provide feedback on platform capabilities, and influence product roadmap decisions.

Partner with customer infrastructure teams to tune, scale, and optimize GPU clusters for cost efficiency, throughput, and AI workload performance.

What we need to see:

5+ years of hands‑on infrastructure or platform engineering experience, with demonstrated expertise designing and operating large-scale GPU clusters (100+ nodes).

Deep expertise building IaaS, PaaS, and SaaS platform layers - architecting and developing infrastructure foundations, not consuming cloud services.

Proficiency in Kubernetes (GPU operator, device plugins, multi-tenancy) and Slurm for HPC and AI workloads.

Hands‑on experience with infrastructure-as-code (Terraform, Helm, Ansible), CI/CD pipelines, and observability stacks (Prometheus, Grafana, DCGC).

Strong coding ability in Python and/or Go/C++, building platform tooling and automation from scratch.

Experience with cloud-native networking (InfiniBand, RoCE, RDMA) and distributed storage solutions for GPU environments.

Excellent communication skills, credibly engaging both infrastructure engineers and C-level stakeholders on complex technical and strategic topics.

Bachelor's degree in Computer Science, Computer Engineering, or equivalent experience.

Ways to stand out from the crowd:

Working knowledge of LLM, MLLM, and RAG frameworks and how they map to infrastructure requirements.

Hands‑on experience with model serving frameworks (Triton Inference Server, vLLM, TensorRT-LLM) and inference optimization techniques.

Proven track record optimizing infrastructure for cost efficiency, throughput, and resource utilization in multi-tenant production environments.

Deep understanding of distributed training concepts (data parallelism, model parallelism, pipeline parallelism) from an infrastructure perspective.

Experience deploying and managing GPU clusters in cloud environments (AWS, Azure, GCP) and on-premises infrastructure at enterprise scale

With competitive salaries and a generous benefits package, we are widely considered to be one of the technology world's most desirable employers. We have some of the most forward-thinking and hardworking people in the world working for us and, due to unprecedented growth, our exclusive engineering teams are rapidly growing. If you're a creative and autonomous engineer with a real passion for technology, we want to hear from you! NVIDIA is committed to fostering a diverse work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

NVIDIA, a pioneer in accelerated computing, revolutionized graphics with the GPU and is at the forefront of AI, gaming, and data-center innovations.

Computer Hardware Manufacturing Computers and Electronics Manufacturing Manufacturing

Company size 10,001+ employees

Company type Public company Founded 1993

Total funding $4B Grant

Momentum

Team growth

Momentum

Team growth

19% in 12 mo

47,502 employees on LinkedIn

Mar 23 May 26

Employee experience

What its like inside

4.3

6,912 reviews

90% would recommend

Culture & values

Culture & values 4.4

Work-life balance

Work-life balance 4.0

Career opportunities

Career opportunities 4.3

Compensation & benefits 4.4

Capital

Funding history

May 2023

Post IPO Equity - ARK Investment Management$65M

Aug 2022

Post IPO Equity

Feb 2021

Post IPO Equity - SoftBank Vision Fund$4B

May 2017

Footprint

Where they work
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Infrastructure Engineer - DGX Cloud, Senior AI Infrastructure Engineer - DGX Cloud
Senior AI Infrastructure Engineer - DGX Cloud, Senior AI Infrastructure Engineer - DGX Cloud

NVIDIA • Bengaluru

On-site
INR 1,500,000 - 2,500,000
Director, Site Reliability and Software Engineering DGX Cloud (Mumbai)
Director, Site Reliability and Software Engineering DGX Cloud (Mumbai)

NVIDIA • Mumbai

On-site
INR 17,143,000 - 26,667,000
Senior Staff Site Reliability Engineer
Senior Staff Site Reliability Engineer

NVIDIA Corporation • India

On-site
INR 4,000,000 - 6,500,000
Senior Solution Architect, Cloud Infrastructure (Maharashtra)
Senior Solution Architect, Cloud Infrastructure (Maharashtra)

NVIDIA • India

On-site
INR 4,000,000 - 7,000,000
Senior Systems Software Engineer, Developer Productivity and Cloud Automation - GeForce NOW
Senior Systems Software Engineer, Developer Productivity and Cloud Automation - GeForce NOW

NVIDIA Corporation • Pune District

On-site
INR 4,000,000 - 6,500,000
Senior Staff Site Reliability Engineer
Senior Staff Site Reliability Engineer

NVIDIA • Bengaluru

On-site
INR 5,000,000 - 7,500,000
Senior Developer Relations Manager - Conglomerates
Senior Developer Relations Manager - Conglomerates

NVIDIA • Mumbai

On-site
INR 4,000,000 - 7,000,000
Senior AI Compute Engineer
Senior AI Compute Engineer

Neysa • Mumbai

On-site
INR 3,500,000 - 6,000,000
Senior Solution Architect, Cloud Infrastructure-DevOps
Senior Solution Architect, Cloud Infrastructure-DevOps

NVIDIA Gruppe • Mumbai

On-site
INR 2,000,000 - 3,000,000
Senior Systems Software Engineer, Developer Productivity and Cloud Automation - GeForce NOW
Senior Systems Software Engineer, Developer Productivity and Cloud Automation - GeForce NOW

NVIDIA Gruppe • Pune District

On-site
INR 400,000 - 660,000
Competitive salary
Generous benefits package