AI/ML Platform Engineer: Kubernetes & GPU Infra at Scale

Deepgram

San Francisco (CA)

On-site

USD 180,000 - 260,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Deepgram is seeking an experienced Platform Engineer to build and operate the hybrid infrastructure foundation for AI/ML research and product development. You will architect, build, and run platforms spanning AWS and on‑premises, enabling teams to train and deploy complex models at scale.

You’ll collaborate with AI researchers and engineers to design tools, automate deployment cycles, and optimize performance and cost across cloud and bare‑metal environments.

Qualifications

  • 5+ years of experience in Platform Engineering, DevOps, or SRE.
  • Hands-on production infrastructure with Terraform.
  • Expert-level Kubernetes knowledge in large-scale environments.
  • Strong scripting and automation skills (e.g., Python, Go, Bash).
  • Experience with CI/CD systems (e.g., GitLab CI, Jenkins, ArgoCD) and building developer tooling.

Responsibilities

  • Architect and maintain the core computing platform using Kubernetes on AWS and on-premises, scalable environment.
  • Develop and manage infrastructure using Terraform as IaC, ensuring reproducibility and automation.
  • Design and optimize AI/ML job scheduling/orchestration with Slurm and Kubernetes integration.
  • Provision, manage, and maintain on-premise bare metal server infrastructure for high-performance GPU computing.
  • Implement networking and storage solutions to support hybrid workloads (CNI, CSI, S3).
  • Develop observability stack (monitoring, logging, tracing) and automation for operations.

Skills

Kubernetes expertise
Terraform
Python
Go
Bash
CI/CD tooling
Automation

Tools

Terraform
GitLab CI/Jenkins/ArgoCD

Job description

Deepgram is seeking an experienced Platform Engineer to build and operate the hybrid infrastructure foundation for AI/ML research and product development. You will architect, build, and run platforms spanning AWS and on‑premises, enabling teams to train and deploy complex models at scale.

You’ll collaborate with AI researchers and engineers to design tools, automate deployment cycles, and optimize performance and cost across cloud and bare‑metal environments.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI/ML Platform Engineer — Hybrid Cloud & GPU
AI/ML Platform Engineer — Hybrid Cloud & GPU

Madrona Venture Labs • United States

Hybrid
USD 180,000 - 260,000
Platform Engineer - AI/ML Infrastructure (Kubernetes & Terraform)
Platform Engineer - AI/ML Infrastructure (Kubernetes & Terraform)

Deepgram • San Francisco (CA)

On-site
USD 180,000 - 260,000
Platform Engineer - AI/ML Infrastructure (Kubernetes & Terraform)
Platform Engineer - AI/ML Infrastructure (Kubernetes & Terraform)

Madrona Venture Labs • United States

On-site
USD 180,000 - 260,000
GPU AI Platform Engineer - Kubernetes & DevOps
GPU AI Platform Engineer - Kubernetes & DevOps

AMD • San Jose (CA)

On-site
USD 140,000 - 170,000
AMD benefits
Lead AI Infrastructure Engineer: Kubernetes & GPU
Lead AI Infrastructure Engineer: Kubernetes & GPU

Seekr • Washington

Hybrid
USD 180,000 - 230,000
Equity ownership
Unlimited PTO
Hybrid work (Reston, VA & Austin, TX)
AI Infrastructure Architect: Kubernetes & GPU Scaling
AI Infrastructure Architect: Kubernetes & GPU Scaling

NVIDIA • United States

Remote
USD 272,000 - 431,000
ML Platform Engineer: Scalable GPU + Kubernetes
ML Platform Engineer: Scalable GPU + Kubernetes

Mistral • Palo Alto (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Healthcare coverage
Parental leave
Relocation support
+2
Senior AI Infra Engineer: GPU Compute on Kubernetes
Senior AI Infra Engineer: GPU Compute on Kubernetes

Harell Data • Palo Alto (CA)

On-site
USD 180,000 - 260,000
Platform Engineer - AI/ML Infrastructure (Kubernetes & Terraform)
Platform Engineer - AI/ML Infrastructure (Kubernetes & Terraform)

Deepgram • United States

On-site
USD 120,000 - 150,000
Medical, dental, vision benefits
Unlimited PTO
Generous paid parental leave
+1
AI Platform Engineer - GPU/Kubernetes Onsite
AI Platform Engineer - GPU/Kubernetes Onsite

AMroute LLC • St. Louis (MO)

On-site
USD 15,429,000 - 26,450,000