Platform Engineer

Harrison Clarke

San Francisco (CA)

On-site

USD 120,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Harrison Clarke is seeking a Senior Platform Engineer to take ownership of their core platform in San Francisco. This position involves designing multi-region Kubernetes clusters, managing GPU infrastructure, overseeing networking systems, and developing observability across metrics and logs. The ideal candidate will have solid experience with Kubernetes in production, infrastructure-as-code tools like Terraform, and must be proficient with GitOps practices. Strong networking fundamentals and familiarity with in-memory data systems are also essential. If you have a passion for building AI systems, apply now!

Qualifications

  • Strong experience operating Kubernetes in production environments, including troubleshooting, autoscaling, and upgrades.
  • Proven background with infrastructure-as-code tools.
  • Hands-on experience running GPU workloads on Kubernetes.

Responsibilities

  • Design and manage multi-region Kubernetes clusters across cloud and GPU-focused providers.
  • Own the deployment lifecycle through GitOps practices.
  • Manage GPU infrastructure, including scheduling efficiency and workload placement.

Skills

Kubernetes
Infrastructure-as-code tools
GPU workloads optimization
GitOps tooling
In-memory data systems
Observability tooling
Networking fundamentals

Tools

Terraform
Helm
Pulumi
Redis

Job description

Overview

Our client, an early-stage company building advanced AI systems, is seeking a senior platform engineer to take ownership of their core platform. This is not a traditional DevOps position focused purely on CI/CD; the role spans GPU orchestration, multi-cloud Kubernetes environments, real-time networking, and observability. The company is already running production workloads across multiple clusters, regions, and hardware types, and is actively expanding into additional cloud providers. This hire will play a key role in scaling and stabilizing that infrastructure.

Key Responsibilities
  • Design and manage multi-region Kubernetes clusters across cloud and GPU-focused providers using infrastructure-as-code
  • Own the deployment lifecycle through GitOps practices (Helm, Kustomize, automated releases, continuous delivery)
  • Manage GPU infrastructure, including scheduling efficiency, workload placement, and cold-start optimization
  • Oversee networking systems such as ingress, gateways, load balancing, and cross-region connectivity
  • Build and maintain observability across metrics, logs, traces, and performance profiling
  • Ensure infrastructure security across identity, secrets, and encryption
  • Maintain CI/CD workflows supporting a monorepo of services and deployment artifacts
  • Partner closely with ML engineers to optimize model serving and GPU utilization
Candidate Profile
  • Strong experience operating Kubernetes in production environments, including troubleshooting, autoscaling, and upgrades
  • Proven background with infrastructure-as-code tools (e.g., Terraform, Pulumi)
  • Hands-on experience running GPU workloads on Kubernetes and understanding resource optimization
  • Familiarity with GitOps tooling such as ArgoCD or Flux, and Helm-based deployments
  • Experience with in-memory data systems (e.g., Redis) and distributed architectures
  • Solid understanding of observability tooling and practices
  • Strong networking fundamentals, particularly in low-latency or distributed systems
  • Experience working in environments with broad ownership across infrastructure
Preferred Background
  • Exposure to GPU cloud providers beyond major hyperscalers
  • Experience with real-time or streaming infrastructure
  • Proficiency in Go or Python
  • Familiarity with ML model deployment and optimization
  • Experience managing infrastructure cost, particularly for GPU-heavy workloads
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GPUaaS Kubernetes Platform Engineer
GPUaaS Kubernetes Platform Engineer

Veriipro • Irving (TX)

On-site
USD 140,000 - 180,000
Senior Kubernetes Engineer
Senior Kubernetes Engineer

GTN Technical Staffing • Dallas (TX)

Hybrid
USD 150,000 - 210,000
Performance bonus
Senior Infrastructure Engineer
Senior Infrastructure Engineer

Harrison Clarke • United States

On-site
USD 100,000 - 140,000
Platform Engineer
Platform Engineer

AMroute LLC • St. Louis (MO)

On-site
USD 15,429,000 - 26,450,000
Member of Technical Staff (Software Engineer, GPU Cluster Infrastructure)
Member of Technical Staff (Software Engineer, GPU Cluster Infrastructure)

United States Digital Space LLC • San Francisco (CA)

On-site
USD 180,000 - 260,000
Software Engineer, AI Infrastructure
Software Engineer, AI Infrastructure

Harell Data • Palo Alto (CA)

On-site
USD 180,000 - 260,000
Sr. Platform Engineer, ML Infrastructure
Sr. Platform Engineer, ML Infrastructure

Insilico Search Partners • Cambridge (MA)

On-site
USD 140,000 - 210,000
Member of Technical Staff (Software Engineer, GPU Cluster Infrastructure)
Member of Technical Staff (Software Engineer, GPU Cluster Infrastructure)

B Capital • United States

On-site
USD 180,000 - 230,000
Senior GPU Infrastructure Engineer
Senior GPU Infrastructure Engineer

Hyperbolic • San Francisco (CA)

On-site
Member of Technical Staff (Software Engineer, GPU Cluster Infrastructure)
Member of Technical Staff (Software Engineer, GPU Cluster Infrastructure)

Perplexity • New York (NY)

On-site
USD 250,000 - 485,000