Client-Facing GPU Kubernetes SRE Engineer

Together Computer Inc

United States

Remote

USD 150,000 - 190,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Together AI seeks an experienced SRE/DevOps engineer to support client-facing Kubernetes GPU clusters in a production environment. The role focuses on high-availability, monitoring, and rapid incident response for AI inference and fine-tuning workloads.

You will work to ensure stability of GPU clusters, collaborate with customers, and guide best practices for deploying and operating AI infrastructure at scale.

Qualifications

  • Experience operating Kubernetes in production with GPU-backed workloads.

Responsibilities

  • Monitor GPU infrastructure health and performance for client clusters.

Skills

Kubernetes
GPU HPC
Slurm
SRE/DevOps
Customer-facing

Tools

Prometheus
Grafana
Terraform
Ansible

Job description

Together AI seeks an experienced SRE/DevOps engineer to support client-facing Kubernetes GPU clusters in a production environment. The role focuses on high-availability, monitoring, and rapid incident response for AI inference and fine-tuning workloads.

You will work to ensure stability of GPU clusters, collaborate with customers, and guide best practices for deploying and operating AI infrastructure at scale.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Technical Support Engineer (GPU Clusters) - US Weekends
Technical Support Engineer (GPU Clusters) - US Weekends

Together Computer Inc • United States

Remote
USD 150,000 - 190,000
AI Infra Engineer – SRE (Kubernetes)
AI Infra Engineer – SRE (Kubernetes)

Berrybytes • United States

On-site
USD 110,000 - 150,000
Principal Kubernetes & AI Infrastructure Engineer
Principal Kubernetes & AI Infrastructure Engineer

NVIDIA • Durham (NC)

On-site
USD 272,000 - 431,000
Equity
Benefits
Platform Engineer - GPU Infra & Kubernetes
Platform Engineer - GPU Infra & Kubernetes

Together AI • San Francisco (CA)

On-site
USD 160,000 - 280,000
Equity
Health insurance
Competitive benefits
Senior Kubernetes & GPU Infra Engineer for AI-scale Compute
Senior Kubernetes & GPU Infra Engineer for AI-scale Compute

Kindredventures • United States

On-site
USD 140,000 - 190,000
GPU Compute Infrastructure Engineer
GPU Compute Infrastructure Engineer

Linuxcareers • San Francisco (CA), Northern (KY)

Hybrid
USD 210,000 - 270,000
AI Infrastructure Engineer: Kubernetes & GPU Clusters
AI Infrastructure Engineer: Kubernetes & GPU Clusters

NVIDIA • United States

Remote
USD 184,000 - 288,000
AI Infrastructure Architect: Kubernetes & GPU Scaling
AI Infrastructure Architect: Kubernetes & GPU Scaling

NVIDIA • United States

Remote
USD 272,000 - 431,000
Kubernetes SRE for AI Infra & GPU Clusters
Kubernetes SRE for AI Infra & GPU Clusters

GMI Cloud • United States

On-site
USD 100,000 - 130,000
Senior AI Infra & Kubernetes Architect
Senior AI Infra & Kubernetes Architect

NVIDIA Corporation • Durham (NC)

On-site
USD 272,000 - 431,000
Equity
Benefits