Platform Reliability Engineer – AI Infra & GPU Fleet (Remote)

Descript

San Francisco (CA)

Hybrid

USD 220,000 - 292,000

Full time

12 days ago
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Catered lunches
Flexible vacation time
401k matching
Remote and hybrid roles

Job summary

Descript seeks an experienced platform engineer to own the foundation of our compute, deployment, reliability, CI/CD, and monorepo health for model training and serving. You will partner with engineers across teams to shape architecture and ensure robust systems.

You will work on on-call processes, tooling, and observability, balancing speed and quality. Descript HQ is in San Francisco, with remote and hybrid roles available.

Qualifications

  • 8+ years building and operating production distributed systems.
  • Experience using agents to multiply impact with AI.
  • Reliability-focused background with incident management experience.
  • Use of SLOs and error budgets as operating tools.
  • Production Kubernetes and IaC in cloud environments.
  • Owned architecture or migration end-to-end with lasting impact.
  • Proactive in identifying unowned work and delivering with minimal guidance.
  • Ability to build minimal repros, read logs, and write focused checks.

Responsibilities

  • Own our platform: GCP, Kubernetes, Temporal, GPU fleet, and deploy/rollback machinery.
  • Own the AI enablement substrate: training/inference pipelines, reliability and cost.
  • Treat cost as an engineering constraint; implement metering and attribution.
  • Ensure security boundaries: IAM, secrets, least privilege, supply-chain integrity.
  • Make infrastructure legible: IaC, runbooks, in-repo docs, observability.
  • Improve how the team learns and ships: hypotheses, instrumentation, incremental releases.
  • Provide architectural direction, mentoring, and clear written communication.

Skills

Distributed systems
Agent-based AI
Reliability engineering
SLOs & error budgets
Kubernetes in production
Infrastructure-as-code
Architecture ownership
Proactive problem solving

Tools

Terraform / IaC
CI/CD pipelines
Cloud platforms (GCP/AWS)
GPU/ML infra tooling

Job description

Descript seeks an experienced platform engineer to own the foundation of our compute, deployment, reliability, CI/CD, and monorepo health for model training and serving. You will partner with engineers across teams to shape architecture and ensure robust systems.

You will work on on-call processes, tooling, and observability, balancing speed and quality. Descript HQ is in San Francisco, with remote and hybrid roles available.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Platform Reliability Engineer – AI-Driven Infra
Platform Reliability Engineer – AI-Driven Infra

Writer • California (MO)

Hybrid
USD 180,000 - 250,000
Generous PTO
Medical, dental, and vision
Parental leave
+4
Platform Reliability Engineer — AI Infra & CloudOps
Platform Reliability Engineer — AI Infra & CloudOps

WRITER • New York (NY)

Hybrid
USD 120,000 - 160,000
Generous PTO
Medical, dental, and vision coverage
Paid parental leave (16 weeks)
+5
Senior AI Infrastructure Platform Engineer
Senior AI Infrastructure Platform Engineer

descript • United States

On-site
USD 220,000 - 292,000
Equity
Competitive benefits
Remote Backend Engineer — Reliability & AI Platform
Remote Backend Engineer — Reliability & AI Platform

Affirm • Chicago (IL)

On-site
USD 173,000 - 233,000
Health care coverage
Flexible Spending Wallets
Time off
+1
SF Platform Engineer — AI Infra, Scale & Reliability
SF Platform Engineer — AI Infra, Scale & Reliability

Harper • San Francisco (CA)

On-site
USD 140,000 - 280,000
Uber commuter benefits
Meals provided (breakfast, lunch, and/
Snacks, drinks and coffee daily
+2
Senior Reliability Engineer — AI Platform
Senior Reliability Engineer — AI Platform

Fireworks AI • San Mateo (CA)

On-site
USD 150,000 - 190,000
Remote Internal Platform Engineer for AI Infra
Remote Internal Platform Engineer for AI Infra

OpenTrain AI • Northern (KY)

Hybrid
USD 120,000 - 155,000
Senior Backend Reliability Engineer — AI‑Driven Platform (Remote)
Senior Backend Reliability Engineer — AI‑Driven Platform (Remote)

Affirm • Riverside (OH)

Remote
USD 173,000 - 233,000
Health coverage
FSA Wallets
Time off
+1
Remote Senior SRE: Build Reliable, Scalable AI Infra
Remote Senior SRE: Build Reliable, Scalable AI Infra

Runware • Town of Sweden (NY)

On-site
USD 140,000 - 190,000
Generous paid time off
Meaningful stock options
Remote-first setup
+3
Remote Backend Engineer: Reliability & AI-Driven Platform
Remote Backend Engineer: Reliability & AI-Driven Platform

Affirm • Denver (CO)

On-site
USD 173,000 - 233,000