Senior DevOps Engineer, Platform Engineering

NVIDIA Gruppe

Santa Clara (CA)

On-site

USD 176,000 - 276,000

Full time

8 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity

Job summary

NVIDIA in Santa Clara, CA seeks a Senior DevOps Platform Engineer for the Metropolis Team to scale CI/CD pipelines and manage Kubernetes-based infrastructure for AI/ML workloads on NVIDIA Data Center GPUs.

You will define release processes, drive developer efficiency with tooling, and own observability using Prometheus, Grafana, and log pipelines. Strong Linux, Python, and Kubernetes skills are required.

Qualifications

  • BS/MS in CS/CE or equivalent experience.
  • Strong Python scripting and automation skills.
  • Kubernetes, Helm, and production orchestration expertise.
  • Experience building scalable CI/CD pipelines at scale.
  • Solid Linux systems administration and networking knowledge.
  • Familiarity with release engineering practices and observability stacks.

Responsibilities

  • Compose, build, and maintain scalable CI/CD pipelines.
  • Develop and manage Kubernetes-based platform infrastructure for AI/ML workloads.
  • Implement scaling and performance measurement frameworks in Kubernetes.
  • Define release engineering processes, branching, versioning, and gating.
  • Drive developer efficiency with tooling and automation frameworks.
  • Own observability and monitoring with Prometheus, Grafana, and logs.

Skills

Python scripting
CI/CD design
Release engineering
Observability

Education

BS or MS in Computer Science/Engineering or equivalent

Tools

Kubernetes
Helm
Jenkins
GitHub Actions
GitLab Actions
Linux
ELK
Prometheus
Grafana

Job description

At NVIDIA, our work is dedicated to a computing model passionate about visual and AI computing. For twenty years, NVIDIA has led the way in visual computing, the science and art of computer graphics, through our invention of the GPU. The GPU has proven extremely effective in solving complex computer science challenges. Today, NVIDIA's GPU powers deep learning algorithms, simulating human intelligence. It serves as the brain for computers, robots, and self‑driving cars that perceive and interpret the world. We aim to expand our company and teams with the brightest minds globally, and now is an exciting time to join us!

Senior DevOps Platform Engineer – Metropolis Team
Responsibilities
  • Compose, build, and maintain scalable CI/CD pipelines using Jenkins, GitHub/GitLab Actions and Runners for Metropolis software products.
  • Develop and manage Kubernetes‑based platform infrastructure supporting AI/ML workloads on NVIDIA Data Center GPUs.
  • Build and implement scaling and performance measurement frameworks within Kubernetes to ensure platform reliability and efficiency under AI/ML workload demands.
  • Define and implement release engineering processes, branching strategies, versioning standards, and gating criteria.
  • Drive developer efficiency by building and maintaining DevOps MCP servers, tooling, and automation frameworks.
  • Own observability and monitoring infrastructure using Prometheus, Grafana, and log aggregation pipelines.
  • Troubleshoot hardware and operating system issues across BareMetal and GPU‑accelerated servers to minimize downtime and maintain platform stability.
Qualifications
  • BS or MS in Computer Science, Computer Engineering, or a related field, or equivalent experience, with over 6+ years of relevant industry background.
  • Advanced skills in Python for scripting, tooling, and automation.
  • Deep expertise with Kubernetes, Helm, and container orchestration in production environments.
  • Verified background in building and maintaining CI/CD pipelines at scale (Jenkins, GitHub/GitLab Actions and Runners, or similar).
  • Solid understanding of Linux systems administration, networking, and distributed systems.
  • Experience with release engineering practices including semantic versioning, release gating, and change management.
  • Hands‑on experience with observability stacks (Prometheus, Grafana, ELK, or similar).
Nice to Have
  • Experience with GPU infrastructure and AI/ML platform engineering at scale.
  • Background in BareMetal and hybrid cloud (AWS, GCP, Azure) environment management.
  • Familiarity with NVIDIA Metropolis, DeepStream, or similar AI video analytics platforms.
  • Experience with GitOps workflows, Infrastructure as Code (Terraform, Ansible).
  • Track record of driving DevOps culture transformation and developer experience improvements.
Compensation and Benefits

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 176,000 USD – 276,000 USD. You will also be eligible for equity and benefits.

Application Deadline

Applications for this job will be accepted until July 14, 2026.

NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. We do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior DevOps Engineer, Platform Engineering
Senior DevOps Engineer, Platform Engineering

NVIDIA AI • Santa Clara (CA)

On-site
USD 176,000 - 276,000
Equity
Benefits
Senior DevOps Engineer, Platform Engineering
Senior DevOps Engineer, Platform Engineering

Socket.dev • Santa Clara (CA)

On-site
USD 176,000 - 276,000
Equity
Comprehensive benefits
Senior DevOps Engineer, Platform Engineering
Senior DevOps Engineer, Platform Engineering

NVIDIA • Santa Clara (CA)

On-site
USD 176,000 - 276,000
Equity
Benefits
Senior DevOps Engineer, Platform Engineering
Senior DevOps Engineer, Platform Engineering

NVIDIA Corporation • Santa Clara (CA), Northern (KY)

Hybrid
USD 176,000 - 276,000
Senior Staff Platform Engineer
Senior Staff Platform Engineer

NVIDIA Corporation • Santa Clara (CA), Northern (KY)

Hybrid
USD 200,000 - 322,000
Senior ML Engineer, Metropolis
Senior ML Engineer, Metropolis

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior DevOps Engineer, AI/ML Platform & CI/CD
Senior DevOps Engineer, AI/ML Platform & CI/CD

NVIDIA • Santa Clara (CA)

On-site
USD 176,000 - 276,000
Equity
Comprehensive benefits
Senior DevOps Platform Engineer — AI/ML Infra & CI/CD
Senior DevOps Platform Engineer — AI/ML Infra & CI/CD

NVIDIA AI • Santa Clara (CA)

On-site
USD 176,000 - 276,000
Equity
Benefits
Senior Systems Software Engineer, Developer Productivity and Cloud Automation - GeForce NOW
Senior Systems Software Engineer, Developer Productivity and Cloud Automation - GeForce NOW

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Comprehensive benefits
Senior Technical Program Manager, NVIDIA Metropolis
Senior Technical Program Manager, NVIDIA Metropolis

NVIDIA AI • Santa Clara (CA)

On-site
USD 168,000 - 322,000
Equity
Benefits