Senior DevOps Platform Engineer - AI/ML at Scale (Equity)

NVIDIA

Santa Clara (CA)

On-site

USD 176,000 - 276,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
Benefits

Job summary

NVIDIA in Santa Clara, CA, seeks a Senior DevOps Platform Engineer to design and scale CI/CD and platform infrastructure for AI/ML workloads on Data Center GPUs.

You will implement release workflows with Jenkins and GitHub/GitLab Actions, manage Kubernetes-based platforms, and build tooling to improve developer productivity and observability using Prometheus, Grafana, and ELK.

This role emphasizes reliability, performance, and cross-team collaboration in a fast-moving AI computing environment.

Qualifications

  • Over 6+ years of relevant industry background.
  • Advanced Python scripting, tooling, and automation.
  • Deep expertise with Kubernetes, Helm, and container orchestration.
  • Experience building and maintaining CI/CD pipelines at scale (Jenkins, GitHub/GitLab Actions and Runners).
  • Solid Linux systems administration, networking, and distributed systems knowledge.
  • Experience with release engineering practices including semantic versioning, release gating, and change management.
  • Hands-on observability stacks (Prometheus, Grafana, ELK).

Responsibilities

  • Build and maintain scalable CI/CD pipelines using Jenkins and GitHub/GitLab Actions.
  • Develop Kubernetes-based platform infrastructure for AI/ML workloads.
  • Implement scaling and performance measurement frameworks in Kubernetes.
  • Define release engineering processes, branching, versioning, gating.
  • Enhance DevOps tooling and automation frameworks.
  • Own observability and monitoring with Prometheus, Grafana, and log pipelines.
  • Troubleshoot hardware and OS across BareMetal and GPU-accelerated servers.

Skills

Python
Kubernetes
CI/CD pipelines
Linux administration
Networking
Observability tools

Education

BS or MS in Computer Science/Engineering
equivalent experience

Tools

Jenkins
GitHub Actions
GitLab Actions
Runners
Prometheus
Grafana
ELK
Terraform
Ansible

Job description

NVIDIA in Santa Clara, CA, seeks a Senior DevOps Platform Engineer to design and scale CI/CD and platform infrastructure for AI/ML workloads on Data Center GPUs.

You will implement release workflows with Jenkins and GitHub/GitLab Actions, manage Kubernetes-based platforms, and build tooling to improve developer productivity and observability using Prometheus, Grafana, and ELK.

This role emphasizes reliability, performance, and cross-team collaboration in a fast-moving AI computing environment.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior DevOps Engineer — AI/ML & Kubernetes (Equity)
Senior DevOps Engineer — AI/ML & Kubernetes (Equity)

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 176,000 - 276,000
Equity
Senior Platform Engineer – AI/ML Infra (Equity)
Senior Platform Engineer – AI/ML Infra (Equity)

NVIDIA Corporation • Santa Clara (CA), Northern (KY)

Hybrid
USD 200,000 - 322,000
Senior DevOps Platform Engineer — AI/ML Infra & CI/CD
Senior DevOps Platform Engineer — AI/ML Infra & CI/CD

NVIDIA AI • Santa Clara (CA)

On-site
USD 176,000 - 276,000
Equity
Benefits
Senior DevOps Engineer, AI/ML Platform & CI/CD
Senior DevOps Engineer, AI/ML Platform & CI/CD

NVIDIA • Santa Clara (CA)

On-site
USD 176,000 - 276,000
Equity
Comprehensive benefits
Senior Full-Stack Engineer — AI Platform, Equity Eligible
Senior Full-Stack Engineer — AI Platform, Equity Eligible

Nvidia Corporation in • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Remote Senior AI Platform Engineer — ML Infrastructure Lead
Remote Senior AI Platform Engineer — ML Infrastructure Lead

NVIDIA Corporation • Santa Clara (CA)

Hybrid
USD 184,000 - 357,000
Equity
Comprehensive benefits package
Senior Platform Services Engineer - Scalable Web & AI Tools
Senior Platform Services Engineer - Scalable Web & AI Tools

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 184,000 - 287,500
Equity
Benefits
Competitive salary
Senior Full-Stack Engineer — AI Infra & GPU Cloud (Equity)
Senior Full-Stack Engineer — AI Infra & GPU Cloud (Equity)

Nvidia Corporation • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Senior DevOps Engineer, Platform Engineering
Senior DevOps Engineer, Platform Engineering

NVIDIA AI • Santa Clara (CA)

On-site
USD 176,000 - 276,000
Equity
Benefits
Senior DevOps Engineer, Platform Engineering
Senior DevOps Engineer, Platform Engineering

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 176,000 - 276,000
Equity