Senior DevOps Platform Engineer — AI/ML Infra & CI/CD

NVIDIA AI

Santa Clara (CA)

On-site

USD 176,000 - 276,000

Full time

5 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
Benefits

Job summary

NVIDIA invites applications for a Senior DevOps Platform Engineer to join the Metropolis team. You will design and manage CI/CD pipelines, Kubernetes-based infrastructure, and release processes for AI/ML workloads on NVIDIA Data Center GPUs.

You will lead tooling, automation, and observability efforts, ensuring scalable, reliable platform operations. Strong scripting in Python and deep Kubernetes expertise are essential for success.

Qualifications

  • BS or MS in Computer Science, Computer Engineering, or a related field, or equivalent experience, with 6+ years of industry background.
  • Advanced skills in Python for scripting, tooling, and automation.
  • Deep expertise with Kubernetes, Helm, and container orchestration in production environments.
  • Proven background in building and maintaining CI/CD pipelines at scale (Jenkins, GitHub/GitLab Actions and Runners, or similar).
  • Solid understanding of Linux systems administration, networking, and distributed systems.
  • Experience with release engineering practices including semantic versioning, release gating, and change management.
  • Hands-on experience with observability stacks (Prometheus, Grafana, ELK, or similar).

Responsibilities

  • Compose, build, and maintain scalable CI/CD pipelines using Jenkins, GitHub/GitLab Actions and Runners for Metropolis software products.
  • Develop and manage Kubernetes-based platform infrastructure supporting AI/ML workloads on NVIDIA Data Center GPUs.
  • Build and implement scaling and performance measurement frameworks within Kubernetes to ensure platform reliability and efficiency under AI/ML workload demands.
  • Define and implement release engineering processes, branching strategies, versioning standards, and gating criteria.
  • Drive developer efficiency by building and maintaining DevOps MCP servers, tooling, and automation frameworks.
  • Own observability and monitoring infrastructure using Prometheus, Grafana, and log aggregation pipelines.
  • Troubleshoot hardware and operating system issues across BareMetal and GPU-accelerated servers to minimize downtime and maintain platform stability.

Skills

Python
Kubernetes
CI/CD pipelines
Linux
Networking
Observability
Release engineering

Education

BS or MS in Computer Science or related field

Tools

Jenkins
GitHub Actions
GitLab Actions
Runners
Helm
Terraform
Ansible
Prometheus
Grafana
ELK stack

Job description

NVIDIA invites applications for a Senior DevOps Platform Engineer to join the Metropolis team. You will design and manage CI/CD pipelines, Kubernetes-based infrastructure, and release processes for AI/ML workloads on NVIDIA Data Center GPUs.

You will lead tooling, automation, and observability efforts, ensuring scalable, reliable platform operations. Strong scripting in Python and deep Kubernetes expertise are essential for success.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior DevOps Engineer, AI/ML Platform & CI/CD
Senior DevOps Engineer, AI/ML Platform & CI/CD

NVIDIA • Santa Clara (CA)

On-site
USD 176,000 - 276,000
Equity
Comprehensive benefits
Senior DevOps Engineer — AI/ML & Kubernetes (Equity)
Senior DevOps Engineer — AI/ML & Kubernetes (Equity)

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 176,000 - 276,000
Equity
Senior DevOps Engineer, Platform Engineering
Senior DevOps Engineer, Platform Engineering

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 176,000 - 276,000
Equity
Senior DevOps Engineer, Platform Engineering
Senior DevOps Engineer, Platform Engineering

NVIDIA • Santa Clara (CA)

On-site
USD 176,000 - 276,000
Equity
Benefits
Senior DevOps Engineer, Platform Engineering
Senior DevOps Engineer, Platform Engineering

Socket.dev • Santa Clara (CA)

On-site
USD 176,000 - 276,000
Equity
Comprehensive benefits
Senior DevOps Platform Engineer - AI/ML at Scale (Equity)
Senior DevOps Platform Engineer - AI/ML at Scale (Equity)

NVIDIA • Santa Clara (CA)

On-site
USD 176,000 - 276,000
Equity
Benefits
Senior DevOps Engineer, Platform Engineering
Senior DevOps Engineer, Platform Engineering

NVIDIA AI • Santa Clara (CA)

On-site
USD 176,000 - 276,000
Equity
Benefits
Senior Platform Architect – AI/ML Infrastructure
Senior Platform Architect – AI/ML Infrastructure

Nvidia Corporation in • Santa Clara (CA)

On-site
USD 200,000 - 322,000
Senior Platform Engineer - AI/ML Infra & Edge CDN
Senior Platform Engineer - AI/ML Infra & Edge CDN

NVIDIA • Santa Clara (CA)

On-site
USD 200,000 - 322,000
Senior Platform Engineer - AI/ML Infra & Global CDN
Senior Platform Engineer - AI/ML Infra & Global CDN

Nvidia Corporation • Santa Clara (CA)

On-site
USD 200,000 - 322,000
Equity
Comprehensive benefits