GPU AI Platform Engineer | Kubernetes & DevOps Leader

Socket.dev

San Jose, Northern (CA, KY)

Hybrid

USD 140,000 - 200,000

Full time

3 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

AMD is seeking a DevOps / Platform Engineer in San Jose, CA to build and operate large-scale GPU compute infrastructure powering AI and ML workloads. The role demands leadership, effective communication, and willingness to contribute across startup-like needs within a large company.

You will design and extend platform capabilities, manage Kubernetes deployments with Helm and GitOps (ArgoCD/Flux), and collaborate with product teams to deliver robust developer platforms and scalable

Qualifications

  • Experience in DevOps, Platform, or Infrastructure Engineering.
  • Deep hands-on experience with Kubernetes and container orchestration at scale.
  • Proven ability to design and deliver platform features that serve internal customers or developer teams.
  • Experience building developer-facing platforms or internal developer portals.
  • Hands-on experience in storage or network engineering within Kubernetes environments (CSI drivers, dynamic provisioning, CNI plugins, or network policy).
  • Experience with Infrastructure as Code tools like Terraform.
  • Background in HPC, Slurm, or GPU-based compute systems for ML/AI workloads.
  • Practical experience with monitoring and observability tools (Prometheus, Grafana, Loki).
  • Understanding of machine learning frameworks (PyTorch, vLLM, SGLang).

Responsibilities

  • Build and extend platform capabilities to enable new workloads.
  • Design and operate scalable orchestration systems using Kubernetes across on-prem and multi-cloud.
  • Develop platform features such as secret management, configuration management, deployment automation.
  • Partner with development teams to extend GPU developer platform with features, APIs, templates, and self-service workflows.
  • Manage service lifecycle within Kubernetes using Helm and GitOps workflows (ArgoCD or Flux).
  • Apply expertise in storage and networking to design CSI drivers, persistent volumes, and network policies.

Skills

Kubernetes
GitOps
ArgoCD
Flux
Terraform
Prometheus
Grafana
Loki
PyTorch
Slurm

Education

Bachelor's or Master's in Computer Science/Engineering

Tools

ArgoCD
Flux
Terraform
Prometheus

Job description

AMD is seeking a DevOps / Platform Engineer in San Jose, CA to build and operate large-scale GPU compute infrastructure powering AI and ML workloads. The role demands leadership, effective communication, and willingness to contribute across startup-like needs within a large company.

You will design and extend platform capabilities, manage Kubernetes deployments with Helm and GitOps (ArgoCD/Flux), and collaborate with product teams to deliver robust developer platforms and scalable

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Infrastructure Engineer
AI Infrastructure Engineer

Advanced Micro Devices • San Jose (CA)

Hybrid
USD 140,000 - 190,000
AI Systems Engineer: HPC & GPU Clusters
AI Systems Engineer: HPC & GPU Clusters

AMD • San Jose (CA)

On-site
USD 180,000 - 260,000
AMD benefits
AI Infrastructure Engineer
AI Infrastructure Engineer

Socket.dev • San Jose (CA), Northern (KY)

Hybrid
USD 140,000 - 200,000
Senior Datacenter Platform Engineer — GPU/AI Infra
Senior Datacenter Platform Engineer — GPU/AI Infra

Advanced Micro Devices, Inc. • Austin (TX)

On-site
USD 110,000 - 160,000
AMD benefits at a glance
Platform Engineer
Platform Engineer

Harrison Clarke • San Francisco (CA)

On-site
USD 120,000 - 160,000
DevOps Platform Engineer: AI CI/CD & Kubernetes
DevOps Platform Engineer: AI CI/CD & Kubernetes

AMD • Longmont (CO)

On-site
USD 140,000 - 190,000
AI Platform Architect - Hyperscale Datacenter Systems
AI Platform Architect - Hyperscale Datacenter Systems

AMD • Austin (TX)

On-site
USD 180,000 - 260,000
AI Systems Engineer: HPC & GPU Clusters
AI Systems Engineer: HPC & GPU Clusters

Advanced Micro Devices, Inc. • San Jose (CA)

On-site
USD 180,000 - 260,000
AI-Driven HPC Systems Design Engineer
AI-Driven HPC Systems Design Engineer

AMD • Austin (TX)

Hybrid
USD 140,000 - 190,000
DevOps Platform Engineer: CI/CD & Kubernetes
DevOps Platform Engineer: CI/CD & Kubernetes

Advanced Micro Devices • Longmont (CO)

Hybrid
USD 120,000 - 160,000
AMD benefits