Senior Solutions Architect, NVIDIA Cloud Partner Operations

NVIDIA

United States

On-site

USD 180,000 - 260,000

Full time

3 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

NVIDIA is seeking a hands-on Solutions Architect to raise Day 2 operations standards across our NVIDIA Cloud Partner ecosystem. You will work with partner engineers to solve real problems, prototype approaches, and leave behind repeatable practices that boost reliability, performance, and economics of AI clouds at scale.

You will lead cross-functional work with product, engineering, and support to drive adoption of NVIDIA platforms, optimize operating models, and ensure day-2 readiness for new

Qualifications

  • BS, MS, or PhD in Computer Science, Electrical/Computer Engineering, Physics, Mathematics, or related field or equivalent.
  • 12+ years in production infrastructure, cloud engineering, SRE, or similar roles; 5+ years in large-scale GPU/AI infra.
  • Hands-on experience with distributed infra under real production load.
  • Deep expertise in at least one Day 2 stack area with hands-on GPU/HPC/cloud work.
  • Experience with Kubernetes/Slurm, GPU scheduling, Prometheus/Grafana/OpenTelemetry, and IaC tooling.
  • Strong Linux knowledge with scripting in Python/Bash for automation.
  • Evidence-based troubleshooting across system boundaries; strong cross-functional collaboration.

Responsibilities

  • Solve hard Day 2 operations problems at scale with partner engineers.
  • Make new technology Day 2 ready and drive adoption in live environments.
  • Improve reliability, performance, and economics using metrics.
  • Raise partner Day 2 maturity across people, process, tooling, and telemetry.
  • Turn validated work into ecosystem capabilities and reference artifacts.
  • Provide field evidence to cross-functional teams to fix issues at the right level.

Skills

Day 2 operations
GPU cloud
Linux
Python
Kubernetes
Terraform/Ansible
Telemetry/Monitoring

Education

BS/MS in CS/EE or related

Tools

DCGM
NCCL
Lustre/GPFS

Job description

NVIDIA is looking for a hands-on Solutions Architect to raise the Day 2 operations bar across our NVIDIA Cloud Partner ecosystem. Day 2 starts when a cluster is installed and validated: keeping the service healthy, adapting it as technology and customer demand change, and improving performance, stability, efficiency and economics over time. You will work with engineers running AI clouds at scale on the problems that decide whether customers stay and whether the next generation of NVIDIA technology lands successfully.

Our job is to work hand in hand with NCPs to solve real problems and drive real optimizations, prove the answer, and turn it into something the next partner can use! This is not an outsourced operations role. The partner owns its cloud; success means leaving its team more capable, not more dependent on ours.

What you'll be doing:
  • Solve hard Day 2 operations problems at scale. Work alongside partner engineers to find the cause, prototype an approach, validate it under representative load, and leave behind a practice their team can operate.
  • Make new technology Day 2 ready. Help partners prepare the operating model for new NVIDIA platforms, capacity, services, and use cases before customers depend on them, and help drive adoption in live environments without degrading service.
  • Improve reliability, performance, and economics together. Use measures such as incident frequency, recovery time, utilization, and cost per token to show where the cloud is losing performance or margin - and whether the fix worked.
  • Raise each partner's Day 2 maturity. Identify and help close the gaps that matter across people, process, tooling, telemetry, security, and incident response.
  • Turn one solution into ecosystem capability. Convert validated work into operating procedures, reference architectures, assessments, automation, and agentic workflows that other NCPs can integrate into their standard operating model.
  • Create the feedback loop only NVIDIA can. Spot patterns across partners early and bring clear field evidence to account teams, support, product, and engineering so repeated problems are fixed at the right level.
What we need to see:

BS, MS, or PhD in Computer Science, Electrical or Computer Engineering, Physics, Mathematics, or a related field - or equivalent experience.

12+ years in production infrastructure, cloud engineering, solutions architecture, site reliability engineering, HPC, or a similar technical role; alternatively, 5+ years of exceptional specialist-level work in large-scale GPU or AI infrastructure.

Experience building, operating, or improving distributed infrastructure under real production load - not only designing or deploying it.

Deep expertise in at least one part of the Day 2 stack, backed by hands-on work with large-scale GPU, HPC, or cloud infrastructure. Relevant technologies may include DCGM, BMC/Redfish, and firmware and driver lifecycle; InfiniBand or high-speed Ethernet, NCCL, and UFM; or high-performance storage such as Lustre, IBM Storage Scale, WEKA, VAST Data, or comparable platforms.

Working experience across the broader operating platform, including Kubernetes or Slurm, GPU scheduling and multi-tenancy, Prometheus, Grafana or OpenTelemetry, and automation with Terraform, Ansible, Argo CD, or similar tooling.

Strong Linux knowledge and enough Python, Bash, or similar experience to automate measurement, diagnosis, validation, or remediation.

A detailed evidence-led approach to troubleshooting across system boundaries, paired with the judgment to make difficult technical findings clear.

The ability to lead sophisticated work with partner engineers and cross-functional teams without direct authority or taking ownership away from the operator.

Strong communication, prioritization, and time-management skills across multiple partner engagements.

Ways to stand out from the crowd:
  • Real world experience operating a GPU cloud, HPC environment, or large-scale AI platform under customer load.
  • Built or matured a 24/7 operations function, including observability, incident and problem management, coverage, and on-call design.
  • Hands on experience with NVIDIA rack-scale platforms such as GB200 or GB300 NVL72 into production, or have hands-on experience with NVIDIA operations technologies such as Spectrum-X, UFM, Base Command Manager, Mission Control, and the GPU or Network Operators.
  • Driven improved fleet health or unit economics through benchmarking, infrastructure as code, GitOps, automated diagnosis, or agent-based remediation.
  • Even if your background doesn't match every line above, we'd love to hear how your experience applies.

With competitive salaries and a generous benefits package, NVIDIA is widely considered to be one of the technology world's most desirable employers. We have some of the most forward-thinking and hardworking people in the world working for us. This role presents an opportunity to have a wide impact at NVIDIA by improving the factory planning funct

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Solutions Architect, NVIDIA Cloud Partner Operations
Senior Solutions Architect, NVIDIA Cloud Partner Operations

NVIDIA • Santa Clara (CA)

On-site
USD 224,000 - 357,000
Senior Solutions Architect, NVIDIA Cloud Partner Operations
Senior Solutions Architect, NVIDIA Cloud Partner Operations

Nvidia Corporation • Santa Clara (CA)

On-site
USD 224,000 - 357,000
Equity
Comprehensive benefits
Solutions Architect - NVIDIA Cloud Partners
Solutions Architect - NVIDIA Cloud Partners

NVIDIA • United States

On-site
USD 184,000 - 357,000
Senior Solutions Architect, NVIDIA Cloud Partners
Senior Solutions Architect, NVIDIA Cloud Partners

NVIDIA • United States

On-site
USD 184,000 - 288,000
Equity
Benefits
Senior Solutions Architect, NVIDIA Cloud Partners
Senior Solutions Architect, NVIDIA Cloud Partners

NVIDIA Corporation • Santa Clara (CA), Northern (KY)

Hybrid
USD 184,000 - 288,000
Equity
Benefits
Senior Solutions Architect, NVIDIA Cloud Partners
Senior Solutions Architect, NVIDIA Cloud Partners

NVIDIA AI • Santa Clara (CA)

On-site
USD 184,000 - 288,000
Solutions Architect - NVIDIA Cloud Partners
Solutions Architect - NVIDIA Cloud Partners

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity compensation
Benefits
Solutions Architect - NVIDIA Cloud Partners
Solutions Architect - NVIDIA Cloud Partners

NVIDIA • Virginia (IL)

On-site
USD 184,000 - 357,000
Senior Manager, NCP and ISV Business Development
Senior Manager, NCP and ISV Business Development

NVIDIA • United States

Hybrid
USD 180,000 - 240,000
NCX Senior Engineer
NCX Senior Engineer

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Benefits