Senior Platform Engineer

STN Incorporated

United States

Hybrid

USD 140,000 - 180,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

STN Incorporated seeks a Senior Platform Engineer to design and operate the multi-tenant platform layer that turns GPU infrastructure into a usable cloud service. This role will own orchestration and automation across Kubernetes, Slurm, and Run:ai to support a scalable GPU cloud offering.

The position reports to the Director, Platform Engineering and requires 6+ years in platform engineering, deep Kubernetes expertise, and strong Go/Python coding skills. Remote US or hybrid in Pleasanton, CA.

Qualifications

  • 6+ years in platform engineering, SRE, or cloud engineering at scale.
  • Deep Kubernetes expertise including CRDs, operators, and multi-tenant patterns.
  • Strong programming skills in Go, Python, or both.
  • Experience operating GPU clusters or AI infrastructure at production scale.
  • Bachelor’s degree in computer science or equivalent experience.

Responsibilities

  • Design and build the orchestration layer (Kubernetes, Slurm, Run:ai, or comparable)
  • Manage multi-tenant isolation including namespaces, networking, storage, and quotas
  • Build customer-facing platform APIs, CLIs, web portals, and SDKs
  • Implement and operate image management, GPU operator, and node provisioning automation
  • Drive infrastructure-as-code and automation across the platform stack
  • Partner with SRE on platform reliability, SLO definition, and observability
  • Support TAM and Support engineers on customer-impacting platform issues
  • Maintain customer environment templates, configuration management, and rollout tooling
  • Participate in architecture review, design discussions, and technical roadmap
  • Drive continuous platform improvement and reduce operational toil

Skills

Go programming
Python programming
SRE / cloud engineering

Education

Bachelor's degree in computer science or equivalent

Tools

Kubernetes
NVIDIA GPU Operator
Slurm
Run:ai
KubeRay
Istio
Linkerd

Job description

Senior Platform Engineer

Platform and software · shared across customers

Reports to: Director, Platform Engineering (or Chief Architect)

Location: Remote (US) or Pleasanton, CA (hybrid)

Department: Cloud Platform Engineering / GPU Platform Engineering

Position summary

The Senior Platform Engineer builds and operates the multi-tenant orchestration, scheduling, and customer-facing platform layer that turns raw GPU infrastructure into a usable cloud service. This role is the software backbone of GPU One (GPUaaS).

Key responsibilities
  • Design and build the orchestration layer (Kubernetes, Slurm, Run:ai, or comparable)

  • Manage multi-tenant isolation including namespaces, networking, storage, and quotas

  • Build customer-facing platform APIs, CLIs, web portals, and SDKs

  • Implement and operate image management, GPU operator, and node provisioning automation

  • Drive infrastructure-as-code and automation across the platform stack

  • Partner with SRE on platform reliability, SLO definition, and observability

  • Support TAM and Support engineers on customer-impacting platform issues

  • Maintain customer environment templates, configuration management, and rollout tooling

  • Participate in architecture review, design discussions, and technical roadmap

  • Drive continuous platform improvement and reduce operational toil

Required qualifications
  • 6+ years in platform engineering, SRE, or cloud engineering at scale

  • Deep Kubernetes expertise including CRDs, operators, and multi-tenant patterns

  • Strong programming skills in Go, Python, or both

  • Experience operating GPU clusters or AI infrastructure at production scale

  • Bachelor’s degree in computer science or equivalent experience

Preferred qualifications
  • Experience with NVIDIA GPU Operator, MIG, MPS, and NCCL operator patterns

  • Familiarity with Slurm operator, Run:ai, KubeRay, or comparable AI orchestration

  • Service mesh experience (Istio, Linkerd) and multi-cluster networking

  • Open source contributions in the cloud-native or AI infrastructure ecosystem

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Platform Engineer — GPU Cloud Orchestration
Senior Platform Engineer — GPU Cloud Orchestration

STN Incorporated • United States

Hybrid
USD 140,000 - 180,000
Platform Engineer
Platform Engineer

Harrison Clarke • San Francisco (CA)

On-site
USD 120,000 - 160,000
Platform Engineer (GPU)
Platform Engineer (GPU)

Vero • United States

On-site
USD 136,000 - 160,000
Medical, dental, and vision insurance
Equity Scheme
401(k) with employer match
+3
Senior Kubernetes Platform Developer – GPU & AI Infrastructure
Senior Kubernetes Platform Developer – GPU & AI Infrastructure

GTN Technical Staffing • Dallas (TX)

Hybrid
USD 165,000 - 210,000
Relocation available
Hybrid work model
Senior Systems and Platform Engineer
Senior Systems and Platform Engineer

Maxisiq • Bethesda (AR)

On-site
USD 180,000 - 260,000
Senior Systems and Platform Engineer
Senior Systems and Platform Engineer

MAXISIQ, Inc. • Bethesda (MD)

On-site
USD 170,000 - 210,000
VP of Engineering
VP of Engineering

Hyperbolic • San Francisco (CA)

On-site
USD 200,000 - 300,000
SRE / Platform Engineer, GPU Infrastructure
SRE / Platform Engineer, GPU Infrastructure

Bake AI • Hillsboro (OR)

On-site
USD 140,000 - 210,000
Senior Kubernetes Developer – GPU & AI Infrastructure
Senior Kubernetes Developer – GPU & AI Infrastructure

GTN Technical Staffing • Town of Texas (WI), Fort Worth (TX)

Hybrid
USD 150,000 - 210,000
Relocation assistance
Hybrid work arrangement
Remote work flexibility
Senior/Staff Software Engineer, Kubernetes Infrastructure
Senior/Staff Software Engineer, Kubernetes Infrastructure

Kindredventures • United States

On-site
USD 140,000 - 190,000