AI GPU Support Engineer — Kubernetes SRE (Remote)

United States Digital Space LLC

United States

Remote

USD 160,000 - 230,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Startup equity
Health insurance
Remote-friendly

Job summary

United States Digital Space LLC is seeking a Technical Support Engineer to be the first line of defense for customers building training, fine tuning, and inference solutions with our AI infrastructure. You will work with Kubernetes GPU clusters, act as an SRE, and serve as a product expert while collaborating with product and sales to improve offerings.

The role requires 3+ years in a customer-facing technical role, deep Kubernetes knowledge, and strong AI/HPC fundamentals.

Qualifications

  • 3+ years in a customer-facing technical role with AI service or mission-critical API support
  • Experience as an SRE or DevOps engineer working with Kubernetes
  • Strong background in AI/ML and HPC environments
  • Advanced knowledge of Kubernetes, SLURM, IaC, high-performance networking, and storage management
  • Experience with HPC/Slurm cluster environments, node draining and job scheduling
  • Familiarity with InfiniBand, RDMA, and network diagnostics
  • Ability to collaborate cross-functionally with Sales, Engineering, Support, Product and Research
  • Excellent communication and ownership mindset

Responsibilities

  • Engage with customers to resolve complex technical challenges in Kubernetes GPU clusters
  • Act as customer-facing SRE to keep GPU clusters healthy and stable
  • Become a product expert for GPU Cluster service and guide escalation to Eng/Prod
  • Monitor cluster health and report hardware issues with remediation steps
  • Maintain production infrastructure for enterprise GPU customers (fleet rebalancing, Slurm maintenance)
  • Investigate storage and networking issues in bare-metal and VM environments
  • Document configurations, procedures, and FAQs for knowledge sharing
  • Provide coverage during holidays, nights and weekends as needed

Skills

Kubernetes
SRE/DevOps
AI/ML basics
HPC environments
IaC (Ansible)
High-performance networking
Troubleshooting complex systems
Cross-functional collaboration

Tools

SLURM
Ansible
NFS

Job description

United States Digital Space LLC is seeking a Technical Support Engineer to be the first line of defense for customers building training, fine tuning, and inference solutions with our AI infrastructure. You will work with Kubernetes GPU clusters, act as an SRE, and serve as a product expert while collaborating with product and sales to improve offerings.

The role requires 3+ years in a customer-facing technical role, deep Kubernetes knowledge, and strong AI/HPC fundamentals.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI GPU Cluster Support Engineer (Remote)
AI GPU Cluster Support Engineer (Remote)

Together AI • United States

Remote
USD 160,000 - 230,000
AI Support Engineer — GPU/Kubernetes Expert (Remote)
AI Support Engineer — GPU/Kubernetes Expert (Remote)

Togetherai • San Francisco (CA)

Hybrid
USD 160,000 - 230,000
Health insurance
Competitive compensation
Startup equity
Remote AI Inference Support Engineer (SRE)
Remote AI Inference Support Engineer (SRE)

Together • United States

Remote
USD 160,000 - 230,000
Startup equity
Health insurance
Remote work flexibility
Remote AI Inference Support Engineer (SRE)
Remote AI Inference Support Engineer (SRE)

Together AI • United States

Remote
USD 160,000 - 230,000
Remote AI Inference Support Engineer
Remote AI Inference Support Engineer

Together AI • New York (NY)

On-site
USD 160,000 - 230,000
Health insurance
Startup equity
Remote work
AI Infrastructure Engineer — GPU Kubernetes for Production
AI Infrastructure Engineer — GPU Kubernetes for Production

vCluster • Germany (OH)

On-site
USD 150,000 - 200,000
Competitive Salary
Platinum-Level Insurance
Flexible Working Schedule
+1
AI Infra Engineer – SRE (Kubernetes)
AI Infra Engineer – SRE (Kubernetes)

Berrybytes • United States

On-site
USD 110,000 - 150,000
AI Infra Engineer: Kubernetes on Bare Metal GPUs (Remote)
AI Infra Engineer: Kubernetes on Bare Metal GPUs (Remote)

vCluster • New York (NY)

Hybrid
USD 150,000 - 200,000
Competitive Salary
Platinum-Level Insurance
Flexible Working Schedule
+1
Remote AI Platform Engineer — Kubernetes & GPU, Equity
Remote AI Platform Engineer — Kubernetes & GPU, Equity

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 250,000 - 300,000
Meaningful equity
Fully remote across North America
Full insurance coverage for you and你的依
Senior Software Engineer, GPU AI Infra on Kubernetes (Equity)
Senior Software Engineer, GPU AI Infra on Kubernetes (Equity)

NVIDIA Gruppe • Seattle (WA)

On-site
USD 184,000 - 357,000
Equity
Benefits package
Performance bonuses