Platform Engineer

Reactor

San Francisco (CA)

On-site

USD 190,000 - 260,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Visa sponsorship
Relocation assistance
Health, dental, vision
Equity

Job summary

Reactor in San Francisco is seeking an infrastructure lead to own the platform our AI models run on. You’ll manage multi-region Kubernetes clusters across cloud providers, implement IaC, and drive GPU orchestration for fast, reliable model serving.

You’ll build a robust observability and networking stack, own security and CI/CD for Go and Python services, and partner with ML engineers to optimize containerized workloads in a high-growth startup.

Qualifications

  • Experience operating Kubernetes at production scale, including node-level debugging and upgrades.
  • Proficient with infrastructure-as-code across environments and regions.
  • Experience with GPU workloads on Kubernetes and GPU-aware scheduling.

Responsibilities

  • Provision and manage multi-region Kubernetes clusters across AWS and GPU clouds using IaC.
  • Own the GitOps deployment lifecycle (Helm, Kustomize, image automation, CD).
  • Manage GPU node infrastructure: scheduling, caching, prefetching, observability.

Skills

Kubernetes at scale
Infrastructure as code
GPU workloads on Kubernetes
GitOps tooling
Helm chart authoring
Observability stack
Networking fundamentals
End-to-end infra ownership
Go or Python proficiency
Container orchestration

Tools

Terraform
Pulumi
FluxCD
ArgoCD
Kustomize
Helm
Redis

Job description

Description

You'll own the infrastructure platform that our AI models run on. This isn't a CI/CD-focused DevOps role. You'll work across GPU orchestration, multi-cloud Kubernetes, real-time networking, and observability. You'll be the person who knows why a model pod took 4 minutes to schedule, why cross-region latency spiked, or why a media relay is dropping packets.

Department

Engineering

Location

San Francisco

We run production today across multiple Kubernetes clusters, regions, and GPU types, and we're actively expanding to additional cloud providers. You'll lead that expansion and keep everything running.

What You'll Do
  • Provision and manage multi-region Kubernetes clusters across AWS and GPU cloud providers using infrastructure-as-code.
  • Own the GitOps deployment lifecycle (Helm charts, Kustomize overlays, image automation, and continuous delivery.)
  • Manage GPU node infrastructure: scheduling, model weight caching, image prefetching for fast cold starts, and GPU observability.
  • Operate and improve our networking layer: ingress and gateway management, load balancing, media relay infrastructure, and cross-region connectivity.
  • Build and maintain our observability stack: metrics, logs, traces, and profiling across all services and GPU workloads.
  • Maintain infrastructure security: IAM, secret management, certificate automation, and encryption at rest.
  • Own CI/CD pipelines for monorepo builds spanning Go services, Python model containers, and Helm chart releases.
  • Partner with ML engineers on model serving: container optimization, health checks and startup tuning, media pipeline performance, and multi-GPU configuration.
What We're Looking For
  • You've operated Kubernetes in production at scale, not just deployed to it, but debugged node-level scheduling issues, tuned autoscalers, and managed cluster upgrades.
  • Strong infrastructure-as-code experience (Terraform, Pulumi, or similar) across multiple environments and regions.
  • You've worked with GPU workloads on Kubernetes: device plugins, node taints/tolerations, GPU-aware scheduling. You understand why bin-packing matters for expensive hardware.
  • Experience with GitOps tooling (FluxCD, ArgoCD, or similar) and Helm chart authoring.
  • Comfortable with Redis or similar in-memory data stores (replication, persistence, pub/sub or streaming patterns)
  • Familiarity with modern observability stacks (Prometheus, Grafana, OpenTelemetry, or equivalent) and knowing when to reach for metrics vs. logs vs. traces.
  • Solid networking fundamentals: load balancers, TLS, DNS, NAT. Real-time or low-latency networking experience is a strong plus.
  • You've worked in a startup where you owned infrastructure end-to-end, not just one slice of it.
Nice to Have
  • Experience with GPU cloud providers beyond AWS (Crusoe, CoreWeave, Lambda Labs, Nebius)
  • Real-time media or streaming infrastructure
  • Go or Python proficiency
  • Familiarity with ML model serving (container image optimization, weight loading, GPU driver and runtime management)
  • FinOps and GPU cost optimization
What We're Not Looking For
  • Pure CI/CD pipeline engineers who haven't operated Kubernetes clusters directly
  • Candidates whose infrastructure experience is limited to managed PaaS (Heroku, Vercel, Railway)
  • People who need a fully defined scope, this role requires figuring out what to build next, not just executing tickets
Benefits
  • Competitive San Francisco salary and meaningful equity
  • We sponsor visas and support relocation to the US
  • Generous health, dental, and vision coverage
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Platform Engineer
Platform Engineer

Harrison Clarke • San Francisco (CA)

On-site
USD 120,000 - 160,000
Senior Infrastructure Engineer
Senior Infrastructure Engineer

Harrison Clarke • United States

On-site
USD 100,000 - 140,000
Member of Technical Staff, DevOps
Member of Technical Staff, DevOps

Reactor • San Francisco (CA)

On-site
USD 100,000 - 160,000
Competitive salary and early equity
Visa sponsorship
Generous health, dental, and vision coverage
Software Engineer, AI Infrastructure
Software Engineer, AI Infrastructure

Harell Data • Palo Alto (CA)

On-site
USD 180,000 - 260,000
Platform Engineer - AI/ML Infrastructure (Kubernetes & Terraform)
Platform Engineer - AI/ML Infrastructure (Kubernetes & Terraform)

Madrona Venture Labs • United States

On-site
USD 180,000 - 260,000
Platform Engineer (GPU)
Platform Engineer (GPU)

Vero • United States

On-site
USD 144,000 - 176,000
Medical, dental, and vision insurance
Equity Scheme
401(k) with employer match
+3
Member of Technical Staff - GPU Infrastructure
Member of Technical Staff - GPU Infrastructure

Hyperbolic • San Francisco (CA)

On-site
USD 180,000 - 260,000
Software Engineer, Compute Foundations
Software Engineer, Compute Foundations

Linuxcareers • San Francisco (CA), Northern (KY)

On-site
USD 210,000 - 270,000
VP of Engineering
VP of Engineering

Hyperbolic • San Francisco (CA)

On-site
USD 200,000 - 300,000
AI Infra Engineer – SRE (Kubernetes)
AI Infra Engineer – SRE (Kubernetes)

Berrybytes • United States

On-site
USD 110,000 - 150,000