AI/ML Platform Engineer — Hybrid Cloud & GPU

Madrona Venture Labs

United States

Hybrid

USD 180,000 - 260,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Deepgram is seeking an experienced Platform Engineer to build and operate the hybrid infrastructure foundation for our AI/ML research and product development. You will architect, build, and run the platform spanning AWS and on‑prem data centers, empowering our teams to train and deploy complex models at scale.

This role emphasizes self‑service environments, Kubernetes, Terraform, and GPU scheduling with Slurm, plus on‑prem bare metal and a focus on observability and automation.

Qualifications

  • 5+ years of Platform Engineering, DevOps, or SRE experience.
  • Hands-on production infrastructure with Terraform.
  • Expert knowledge of Kubernetes architecture and operations at scale.
  • Strong scripting and automation (Python, Go, Bash).
  • Experience with CI/CD systems and developer tooling.

Responsibilities

  • Architect and maintain our core computing platform using Kubernetes on AWS and on‑premise.
  • Develop and manage infrastructure using Terraform to ensure reproducibility and automation.
  • Design and optimize AI/ML job scheduling and orchestration with Slurm and Kubernetes.
  • Provision, manage, and maintain on‑premise bare metal server infrastructure for GPU computing.
  • Implement networking and storage to support hybrid workloads.
  • Develop observability, monitoring, logging, and automation for operations.
  • Collaborate with AI researchers to build tools and workflows that accelerate development.

Skills

Kubernetes
Terraform
Slurm
AWS
Python
Go
Bash
CI/CD tooling

Tools

Terraform
Kubernetes
Slurm
AWS
GitLab CI
Jenkins
ArgoCD

Job description

Deepgram is seeking an experienced Platform Engineer to build and operate the hybrid infrastructure foundation for our AI/ML research and product development. You will architect, build, and run the platform spanning AWS and on‑prem data centers, empowering our teams to train and deploy complex models at scale.

This role emphasizes self‑service environments, Kubernetes, Terraform, and GPU scheduling with Slurm, plus on‑prem bare metal and a focus on observability and automation.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Platform Engineer - AI/ML Infrastructure (Kubernetes & Terraform)
Platform Engineer - AI/ML Infrastructure (Kubernetes & Terraform)

Madrona Venture Labs • United States

Hybrid
USD 180,000 - 260,000
Platform Engineer - AI/ML Infrastructure (Kubernetes & Terraform)
Platform Engineer - AI/ML Infrastructure (Kubernetes & Terraform)

Deepgram • United States

Hybrid
USD 120,000 - 150,000
Medical, dental, vision benefits
Unlimited PTO
Generous paid parental leave
+1
Hybrid AI HPC Infrastructure Engineer (GPU/ML)
Hybrid AI HPC Infrastructure Engineer (GPU/ML)

Analysis Group, Inc. • Boston (MA)

On-site
USD 150,000 - 170,000
Discretionary annual bonus
Benefits package
Senior Cloud Infra Engineer (Hybrid) – AI & GPU
Senior Cloud Infra Engineer (Hybrid) – AI & GPU

Colossus Technologies Group • Boston (MA)

Hybrid
USD 180,000 - 220,000
Health & wellness benefit
Competitive equity package
Hybrid work flexibility
Senior AI Infrastructure & Tooling Programs Lead
Senior AI Infrastructure & Tooling Programs Lead

Apply • United States

On-site
USD 150,000 - 190,000
Senior Cloud Platform Engineer - GPU Infra, Hybrid
Senior Cloud Platform Engineer - GPU Infra, Hybrid

Neura Market • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Wellness stipend
Commuter stipend
401k with company match
Senior AI Platform PM: Real-Time ML Infrastructure
Senior AI Platform PM: Real-Time ML Infrastructure

AI Chopping Block • United States

On-site
USD 180,000 - 240,000
Staff AI Platform Engineer — Scalable ML Infra & GPUs
Staff AI Platform Engineer — Scalable ML Infra & GPUs

Jobzhr • Mountain View (CA)

Hybrid
USD 175,000 - 287,000
Platform Engineer: AI Infra for Hybrid Cloud & Kubernetes
Platform Engineer: AI Infra for Hybrid Cloud & Kubernetes

Meibel • Washington

On-site
USD 120,000 - 180,000
Senior ML Platform Engineer — GPU, Kubernetes Infra
Senior ML Platform Engineer — GPU, Kubernetes Infra

WorkGenius Group • Los Angeles (CA)

On-site
USD 117,000 - 186,000