Platform Engineer

Harrison Clarke

San Francisco (CA)

On-site

USD 120,000 - 160,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Harrison Clarke is seeking a Senior Platform Engineer to take ownership of their core platform in San Francisco. This position involves designing multi-region Kubernetes clusters, managing GPU infrastructure, overseeing networking systems, and developing observability across metrics and logs. The ideal candidate will have solid experience with Kubernetes in production, infrastructure-as-code tools like Terraform, and must be proficient with GitOps practices. Strong networking fundamentals and familiarity with in-memory data systems are also essential. If you have a passion for building AI systems, apply now!

Qualifications

  • Strong experience operating Kubernetes in production environments, including troubleshooting, autoscaling, and upgrades.
  • Proven background with infrastructure-as-code tools.
  • Hands-on experience running GPU workloads on Kubernetes.

Responsibilities

  • Design and manage multi-region Kubernetes clusters across cloud and GPU-focused providers.
  • Own the deployment lifecycle through GitOps practices.
  • Manage GPU infrastructure, including scheduling efficiency and workload placement.

Skills

Kubernetes
Infrastructure-as-code tools
GPU workloads optimization
GitOps tooling
In-memory data systems
Observability tooling
Networking fundamentals

Tools

Terraform
Helm
Pulumi
Redis

Job description

Overview

Our client, an early-stage company building advanced AI systems, is seeking a senior platform engineer to take ownership of their core platform. This is not a traditional DevOps position focused purely on CI/CD; the role spans GPU orchestration, multi-cloud Kubernetes environments, real-time networking, and observability. The company is already running production workloads across multiple clusters, regions, and hardware types, and is actively expanding into additional cloud providers. This hire will play a key role in scaling and stabilizing that infrastructure.

Key Responsibilities
  • Design and manage multi-region Kubernetes clusters across cloud and GPU-focused providers using infrastructure-as-code
  • Own the deployment lifecycle through GitOps practices (Helm, Kustomize, automated releases, continuous delivery)
  • Manage GPU infrastructure, including scheduling efficiency, workload placement, and cold-start optimization
  • Oversee networking systems such as ingress, gateways, load balancing, and cross-region connectivity
  • Build and maintain observability across metrics, logs, traces, and performance profiling
  • Ensure infrastructure security across identity, secrets, and encryption
  • Maintain CI/CD workflows supporting a monorepo of services and deployment artifacts
  • Partner closely with ML engineers to optimize model serving and GPU utilization
Candidate Profile
  • Strong experience operating Kubernetes in production environments, including troubleshooting, autoscaling, and upgrades
  • Proven background with infrastructure-as-code tools (e.g., Terraform, Pulumi)
  • Hands-on experience running GPU workloads on Kubernetes and understanding resource optimization
  • Familiarity with GitOps tooling such as ArgoCD or Flux, and Helm-based deployments
  • Experience with in-memory data systems (e.g., Redis) and distributed architectures
  • Solid understanding of observability tooling and practices
  • Strong networking fundamentals, particularly in low-latency or distributed systems
  • Experience working in environments with broad ownership across infrastructure
Preferred Background
  • Exposure to GPU cloud providers beyond major hyperscalers
  • Experience with real-time or streaming infrastructure
  • Proficiency in Go or Python
  • Familiarity with ML model deployment and optimization
  • Experience managing infrastructure cost, particularly for GPU-heavy workloads
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Platform Engineer
Platform Engineer

Reactor • San Francisco (CA)

On-site
USD 190,000 - 260,000
Visa sponsorship
Relocation assistance
Health, dental, vision
+1
Platform Architect - HPC, Kubernetes
Platform Architect - HPC, Kubernetes

EPAM Systems Inc • United States

Remote
USD 140,000 - 230,000
Senior Infrastructure Engineer
Senior Infrastructure Engineer

Harrison Clarke • United States

On-site
USD 100,000 - 140,000
Software Engineer, AI Infrastructure
Software Engineer, AI Infrastructure

Harell Data • Palo Alto (CA)

On-site
USD 180,000 - 260,000
Senior Platform Engineer
Senior Platform Engineer

STN Incorporated • United States

Hybrid
USD 140,000 - 180,000
Senior Platform Engineer — GPU Cloud Infrastructure
Senior Platform Engineer — GPU Cloud Infrastructure

StratITech • San Francisco (CA)

On-site
USD 210,000 - 260,000
Senior AI Infrastructure Engineer, Kubernetes
Senior AI Infrastructure Engineer, Kubernetes

Engg • San Francisco (CA)

On-site
USD 180,000 - 240,000
Platform Engineer - AI/ML Infrastructure (Kubernetes & Terraform)
Platform Engineer - AI/ML Infrastructure (Kubernetes & Terraform)

Madrona Venture Labs • United States

On-site
USD 180,000 - 260,000
Member of Technical Staff - GPU Infrastructure
Member of Technical Staff - GPU Infrastructure

Hyperbolic • San Francisco (CA)

On-site
USD 180,000 - 260,000
Software Engineer, Compute Foundations
Software Engineer, Compute Foundations

Linuxcareers • San Francisco (CA), Northern (KY)

On-site
USD 210,000 - 270,000