Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.
NVIDIA is seeking a hands-on Solutions Architect to raise Day 2 operations standards across our NVIDIA Cloud Partner ecosystem. You will work with partner engineers to solve real problems, prototype approaches, and leave behind repeatable practices that boost reliability, performance, and economics of AI clouds at scale.
You will lead cross-functional work with product, engineering, and support to drive adoption of NVIDIA platforms, optimize operating models, and ensure day-2 readiness for new
NVIDIA is looking for a hands-on Solutions Architect to raise the Day 2 operations bar across our NVIDIA Cloud Partner ecosystem. Day 2 starts when a cluster is installed and validated: keeping the service healthy, adapting it as technology and customer demand change, and improving performance, stability, efficiency and economics over time. You will work with engineers running AI clouds at scale on the problems that decide whether customers stay and whether the next generation of NVIDIA technology lands successfully.
Our job is to work hand in hand with NCPs to solve real problems and drive real optimizations, prove the answer, and turn it into something the next partner can use! This is not an outsourced operations role. The partner owns its cloud; success means leaving its team more capable, not more dependent on ours.
BS, MS, or PhD in Computer Science, Electrical or Computer Engineering, Physics, Mathematics, or a related field - or equivalent experience.
12+ years in production infrastructure, cloud engineering, solutions architecture, site reliability engineering, HPC, or a similar technical role; alternatively, 5+ years of exceptional specialist-level work in large-scale GPU or AI infrastructure.
Experience building, operating, or improving distributed infrastructure under real production load - not only designing or deploying it.
Deep expertise in at least one part of the Day 2 stack, backed by hands-on work with large-scale GPU, HPC, or cloud infrastructure. Relevant technologies may include DCGM, BMC/Redfish, and firmware and driver lifecycle; InfiniBand or high-speed Ethernet, NCCL, and UFM; or high-performance storage such as Lustre, IBM Storage Scale, WEKA, VAST Data, or comparable platforms.
Working experience across the broader operating platform, including Kubernetes or Slurm, GPU scheduling and multi-tenancy, Prometheus, Grafana or OpenTelemetry, and automation with Terraform, Ansible, Argo CD, or similar tooling.
Strong Linux knowledge and enough Python, Bash, or similar experience to automate measurement, diagnosis, validation, or remediation.
A detailed evidence-led approach to troubleshooting across system boundaries, paired with the judgment to make difficult technical findings clear.
The ability to lead sophisticated work with partner engineers and cross-functional teams without direct authority or taking ownership away from the operator.
Strong communication, prioritization, and time-management skills across multiple partner engagements.
With competitive salaries and a generous benefits package, NVIDIA is widely considered to be one of the technology world's most desirable employers. We have some of the most forward-thinking and hardworking people in the world working for us. This role presents an opportunity to have a wide impact at NVIDIA by improving the factory planning funct