Senior DevOps Engineer, Platform Engineering

NVIDIA Corporation

Santa Clara, Northern (CA, KY)

On-site

USD 176,000 - 276,000

Full time

8 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

NVIDIA Corporation in Santa Clara is seeking a Senior DevOps Platform Engineer to design and maintain scalable CI/CD pipelines and Kubernetes-based infrastructure for AI/ML workloads on Data Center GPUs.

You will drive release engineering, observability, and automation while supporting BareMetal and GPU-accelerated servers to ensure platform reliability at scale.

Qualifications

  • BS or MS in Computer Science, Computer Engineering, or related field.
  • 6+ years of relevant industry background.
  • Strong Python scripting skills and automation experience.
  • Deep Kubernetes/Helm expertise in production.
  • CI/CD pipelines experience (Jenkins, GitHub/GitLab Actions).
  • Strong Linux systems administration, networking, and distributed systems knowledge.
  • Hands-on observability experience (Prometheus, Grafana, ELK).

Responsibilities

  • Compose, build, and maintain scalable CI/CD pipelines using Jenkins, GitHub/GitLab Actions and Runners for Metropolis software products.
  • Develop and manage Kubernetes-based platform infrastructure supporting AI/ML workloads on NVIDIA Data Center GPUs.
  • Build scaling and performance measurement frameworks within Kubernetes for reliability under AI/ML workloads.
  • Define and implement release engineering processes, branching strategies, versioning standards, and gating criteria.
  • Drive developer efficiency by building and maintaining DevOps MCP servers, tooling, and automation frameworks.
  • Own observability and monitoring infrastructure using Prometheus, Grafana, and log aggregation pipelines.
  • Troubleshoot hardware and operating system issues across BareMetal and GPU-accelerated servers to minimize downtime.

Skills

Python scripting
Linux administration
Distributed systems
Observability concepts

Education

BS or MS in Computer Science/Engineering or related field

Tools

Kubernetes
Helm
Jenkins
GitHub Actions
GitLab CI
Prometheus
Grafana
ELK Stack
Terraform
Ansible

Job description

## Senior DevOps Engineer, Platform EngineeringApplylocations: US, CA, Santa Claratime type: Full timeposted on: Posted Yesterdayjob requisition id: JR1968596At NVIDIA, our work is dedicated to a computing model passionate about visual and AI computing. For twenty years, NVIDIA has led the way in visual computing, the science and art of computer graphics, through our invention of the GPU. The GPU has proven extremely effective in solving complex computer science challenges. Today, NVIDIA's GPU powers deep learning algorithms, simulating human intelligence. It serves as the brain for computers, robots, and self-driving cars that perceive and interpret the world. We aim to expand our company and teams with the brightest minds globally, and now is an exciting time to join us!NVIDIA invites applications for a Senior DevOps Platform Engineer skilled in Platform and Release Engineering to join the Metropolis team. The role involves developing, building, and maintaining foundational infrastructure and CI/CD systems that run AI/Machine Learning video analytics workloads at scale using NVIDIA Data Center GPUs. You will foster engineering rigor by setting up reliable release workflows, automation systems, and developer tools to improve efficiency on the Metropolis platform.**What you'll be doing:*** Compose, build, and maintain scalable CI/CD pipelines using Jenkins, GitHub/GitLab Actions and Runners for Metropolis software products.* Develop and manage Kubernetes-based platform infrastructure supporting AI/ML workloads on NVIDIA Data Center GPUs.* Build and implement scaling and performance measurement frameworks within Kubernetes to ensure platform reliability and efficiency under AI/ML workload demands.* Define and implement release engineering processes, branching strategies, versioning standards, and gating criteria.* Drive developer efficiency by building and maintaining DevOps MCP servers, tooling, and automation frameworks.* Own observability and monitoring infrastructure using Prometheus, Grafana, and log aggregation pipelines.* Troubleshoot hardware and operating system issues across BareMetal and GPU-accelerated servers to minimize downtime and maintain platform stability.**What we need to see:*** BS or MS in Computer Science, Computer Engineering, or a related field, or equivalent experience, with over 6+ years of relevant industry background.* Advanced skills in Python for scripting, tooling, and automation.* Deep expertise with Kubernetes, Helm, and container orchestration in production environments.* Verified background in building and maintaining CI/CD pipelines at scale (Jenkins, GitHub/GitLab Actions and Runners, or similar).* Solid understanding of Linux systems administration, networking, and distributed systems.* Experience with release engineering practices including semantic versioning, release gating, and change management.* Hands-on experience with observability stacks (Prometheus, Grafana, ELK, or similar).**Ways to stand out from the crowd:*** Experience with GPU infrastructure and AI/ML platform engineering at scale.* Background in BareMetal and hybrid cloud (AWS, GCP, Azure) environment management.* Familiarity with NVIDIA Metropolis, DeepStream, or similar AI video analytics platforms.* Experience with GitOps workflows, Infrastructure as Code (Terraform, Ansible).* Track record of driving DevOps culture transformation and developer experience improvements.NVIDIA is often viewed as one of the most attractive companies to work for in the technology world. We have some of the most progressive and committed professionals in the field on our team. If you are inventive and independent, we want to hear from you!Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 176,000 USD - 276,000 USD.You will also be eligible for equity and benefits.Applications for this job will be accepted at least until August 17, 2026.This posting is for an existing vacancy.NVIDIA uses AI tools in its recruiting processes.NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior DevOps Engineer, Platform Engineering
Senior DevOps Engineer, Platform Engineering

NVIDIA AI • Santa Clara (CA)

On-site
USD 176,000 - 276,000
Equity
Benefits
Senior DevOps Engineer, Platform Engineering
Senior DevOps Engineer, Platform Engineering

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 176,000 - 276,000
Equity
Senior DevOps Engineer, Platform Engineering
Senior DevOps Engineer, Platform Engineering

NVIDIA • Santa Clara (CA)

On-site
USD 176,000 - 276,000
Equity
Benefits
Senior Platform AI Engineer
Senior Platform AI Engineer

NVIDIA Corporation • Santa Clara (CA)

Hybrid
USD 184,000 - 357,000
Equity
Comprehensive benefits package
Senior DevOps Engineer, Platform Engineering
Senior DevOps Engineer, Platform Engineering

Socket.dev • Santa Clara (CA)

On-site
USD 176,000 - 276,000
Equity
Comprehensive benefits
Senior Software Engineer, At Scale Compute Analysis
Senior Software Engineer, At Scale Compute Analysis

NVIDIA Corporation • Santa Clara (CA)

Hybrid
USD 152,000 - 242,000
Health coverage
Dental + Vision
401(k) match
+6
Senior Platform Telemetry Engineer
Senior Platform Telemetry Engineer

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 152,000 - 288,000
Senior Platform and EngOps Engineer - Cluster Operations
Senior Platform and EngOps Engineer - Cluster Operations

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 176,000 - 276,000
Equity options
Comprehensive benefits package
Senior Full-Stack Lead Engineer
Senior Full-Stack Lead Engineer

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 224,000 - 357,000
Senior Software Engineer, Infrastructure Automation and Distributed Systems
Senior Software Engineer, Infrastructure Automation and Distributed Systems

NVIDIA Corporation • United States

Hybrid
USD 224,000 - 431,250
Equity
Benefits