AI GPU Cluster Support Engineer (Remote)

Together AI

United States

Remote

USD 160,000 - 230,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Together AI is seeking a Technical Support Engineer to be the first line of defense for customers building training, fine tuning, and inference solutions on our Kubernetes GPU cluster platform.

You will work in the Customer Experience organization, collaborating with product, sales, and engineering to improve our offerings and customer success. This role requires shifting to a 4-day weekend schedule after ramp-up and includes on-call coverage on weekends.

Qualifications

  • 3+ years in a customer-facing technical role in AI/SaaS.
  • Experience as an SRE or DevOps with Kubernetes.
  • Strong AI/ML and GPU HPC knowledge.
  • Infrastructure as code (e.g., Ansible) and Slurm experience.
  • Familiar with InfiniBand, RDMA, and high-speed networks.
  • Experience with HPC/Slurm clusters.
  • Ability to collaborate cross-functionally.
  • Excellent communication skills.

Responsibilities

  • Engage directly with customers to tackle and resolve complex technical challenges involving our cutting-edge Kubernetes GPU clusters; ensure swift and effective solutions every time.
  • Act as a customer facing SRE to ensure our customer’s Kubernetes clusters remain healthy and stable
  • Become a product expert in our GPU Cluster service, serving as the last line of technical defense before issues are escalated to Engineering and Product teams.
  • Monitor GPU cluster health and proactively communicate hardware issues to customers with clear remediation steps
  • Operate and maintain production infrastructure for enterprise GPU customers, including fleet rebalancing, Slurm cluster maintenance, node repair/migration, and Kubernetes-based workload management
  • Investigate and resolve storage and networking issues such as Weka filesystem degradation, InfiniBand link failures, and bandwidth anomalies on bare-metal and VM environments
  • Collaborate across Engineering, Research, and Product teams to address customer concerns; collaborate with senior leaders to ensure customer satisfaction.
  • Transform customer insights into action by identifying patterns in support cases and work with Engineering and Go-To-Market teams to drive roadmap.
  • Maintain detailed documentation of system configurations, procedures, troubleshooting guides, and FAQs.
  • Be flexible in providing support coverage during holidays, nights and weekends as required by business needs.

Skills

Kubernetes
SRE/DevOps
AI/ML HPC
Infrastructure as Code
Distributed storage
InfiniBand networking
Networking basics
GPU HPC tooling

Tools

Ansible
Slurm
Weka
NFS

Job description

Together AI is seeking a Technical Support Engineer to be the first line of defense for customers building training, fine tuning, and inference solutions on our Kubernetes GPU cluster platform.

You will work in the Customer Experience organization, collaborating with product, sales, and engineering to improve our offerings and customer success. This role requires shifting to a 4-day weekend schedule after ramp-up and includes on-call coverage on weekends.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI GPU Support Engineer — Kubernetes SRE (Remote)
AI GPU Support Engineer — Kubernetes SRE (Remote)

United States Digital Space LLC • United States

Remote
USD 160,000 - 230,000
Startup equity
Health insurance
Remote-friendly
AI Support Engineer — GPU/Kubernetes Expert (Remote)
AI Support Engineer — GPU/Kubernetes Expert (Remote)

Togetherai • San Francisco (CA)

Hybrid
USD 160,000 - 230,000
Health insurance
Competitive compensation
Startup equity
Remote AI Inference Support Engineer
Remote AI Inference Support Engineer

Together AI • New York (NY)

On-site
USD 160,000 - 230,000
Health insurance
Startup equity
Remote work
AI Infrastructure Support Engineer
AI Infrastructure Support Engineer

Neura Market • Northern (KY)

Hybrid
USD 160,000 - 230,000
Health insurance
Remote work flexibility
Equity
Customer Support Engineer (GPU Cluster)
Customer Support Engineer (GPU Cluster)

Togetherai • San Francisco (CA)

Hybrid
USD 160,000 - 230,000
Health insurance
Competitive compensation
Startup equity
Remote AI Inference Support Engineer (SRE)
Remote AI Inference Support Engineer (SRE)

Together • United States

Remote
USD 160,000 - 230,000
Startup equity
Health insurance
Remote work flexibility
Remote AI Inference Support Engineer (SRE)
Remote AI Inference Support Engineer (SRE)

Together AI • United States

Remote
USD 160,000 - 230,000
Senior AI Support Engineer — Gen AI & GPU Expert (Remote)
Senior AI Support Engineer — Gen AI & GPU Expert (Remote)

Togetherai • San Francisco (CA)

Hybrid
USD 160,000 - 230,000
Competitive compensation
Startup equity
Health insurance
+1
AI Infrastructure Engineer — GPU Kubernetes for Production
AI Infrastructure Engineer — GPU Kubernetes for Production

vCluster • Germany (OH)

On-site
USD 150,000 - 200,000
Competitive Salary
Platinum-Level Insurance
Flexible Working Schedule
+1
Technical Support Engineer (Inference) - US Weekends
Technical Support Engineer (Inference) - US Weekends

Together • United States

Remote
USD 160,000 - 230,000
Startup equity
Health insurance
Remote work flexibility