Technical Support Engineer (GPU Clusters) - US Weekends

Together Computer Inc

United States

Remote

USD 150,000 - 190,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Together AI seeks an experienced SRE/DevOps engineer to support client-facing Kubernetes GPU clusters in a production environment. The role focuses on high-availability, monitoring, and rapid incident response for AI inference and fine-tuning workloads.

You will work to ensure stability of GPU clusters, collaborate with customers, and guide best practices for deploying and operating AI infrastructure at scale.

Qualifications

  • Experience operating Kubernetes in production with GPU-backed workloads.

Responsibilities

  • Monitor GPU infrastructure health and performance for client clusters.

Skills

Kubernetes
GPU HPC
Slurm
SRE/DevOps
Customer-facing

Tools

Prometheus
Grafana
Terraform
Ansible

Job description

Together AI provides AI infrastructure, and this role focuses on supporting client-facing Kubernetes GPU clusters. Ideal for an SRE or DevOps engineer with deep experience in GPU technologies and HPC environments, who thrives on solving complex technical issues for customers.

This role involves providing technical support for Together AI's GPU clusters, specifically focusing on customer-facing SRE responsibilities. The core work entails resolving complex technical issues that arise within client Kubernetes GPU environments. This includes maintaining cluster stability and ensuring the reliable operation of AI infrastructure used for inference and fine-tuning.

A key aspect of this position is proactive monitoring of the GPU infrastructure's health. When hardware incidents occur, the engineer is responsible for reporting these to clients. The role also encompasses supporting production infrastructure, which includes managing Slurm, performing node repairs, handling migrations, and ensuring the smooth operation of Kubernetes workloads.

Success in this role means ensuring high availability and performance of client GPU clusters, minimizing downtime through effective problem-solving, and providing expert technical guidance. The engineer acts as a critical interface between Together AI's infrastructure and its customers, directly impacting client satisfaction and operational efficiency of their AI/ML workloads.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Client-Facing GPU Kubernetes SRE Engineer
Client-Facing GPU Kubernetes SRE Engineer

Together Computer Inc • United States

Remote
USD 150,000 - 190,000
Customer Support Engineer (GPU Cluster), India
Customer Support Engineer (GPU Cluster), India

Together Computer Inc • United States

Hybrid
USD 90,000 - 150,000
Startup equity
Health insurance
Remote work flexibility
AI GPU Cluster Support Engineer
AI GPU Cluster Support Engineer

Together Computer Inc • United States

Hybrid
USD 90,000 - 150,000
Startup equity
Health insurance
Remote work flexibility
Customer Success Engineer (CSE), GPU Cluster
Customer Success Engineer (CSE), GPU Cluster

Together AI • San Francisco (CA)

On-site
USD 260,000 - 290,000
Health insurance
Startup equity
Flexible remote work
AI Infra Engineer – SRE (Kubernetes)
AI Infra Engineer – SRE (Kubernetes)

Berrybytes • United States

On-site
USD 110,000 - 150,000
Senior GPU Compute Solutions Engineer
Senior GPU Compute Solutions Engineer

NVIDIA • Santa Clara (CA)

On-site
USD 140,000 - 224,000
Equity
Benefits
Customer Support Engineer – GPU/AI Infra
Customer Support Engineer – GPU/AI Infra

The San Francisco Compute Company • San Francisco (CA)

On-site
USD 90,000 - 120,000
Generous equity grant
Visa sponsorships
Retirement matching
+5
AI Infrastructure Engineer: Kubernetes & GPU Clusters
AI Infrastructure Engineer: Kubernetes & GPU Clusters

NVIDIA • United States

Remote
USD 184,000 - 288,000
Senior GPU Cluster Infra Engineer | Remote
Senior GPU Cluster Infra Engineer | Remote

AISafety • Berkeley (CA)

Hybrid
USD 120,000 - 180,000
Health Insurance
401(k) match
PTO 25 days per year
+3
Senior GPU Infrastructure Engineer — HPC & Clusters
Senior GPU Infrastructure Engineer — HPC & Clusters

Prime Intellect AI • San Francisco (CA)

On-site
USD 150,000 - 300,000