Technical Account Manager

asobbi

United States

Remote

USD 120,000 - 150,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A high-growth GPU infrastructure provider is seeking a hands-on Technical Account Manager to support customers running AI workloads on large-scale GPU clusters. This role demands strong Kubernetes experience for troubleshooting and leading incident management. Candidates should have over 7 years in customer-facing roles with expertise in GPU environments. You'll optimize performance and reliability while providing technical leadership and documentation. This is not a ticket-routing position but an engineering-led role.

Qualifications

  • 7+ years in customer-facing infrastructure roles.
  • Strong hands-on Kubernetes experience.
  • Experience operating bare metal GPU clusters.
  • Linux systems and networking fundamentals.
  • Familiarity with GPU drivers and CUDA stack.

Responsibilities

  • Act as the primary technical contact for a portfolio of customers.
  • Troubleshoot production issues across GPU infrastructure.
  • Lead incident triage and mitigation.
  • Support distributed training and inference workloads.
  • Help customers optimize performance and reliability.
  • Create documentation and best practices.
  • Run regular technical check-ins.

Skills

Customer-facing infrastructure support
Kubernetes expertise
Linux systems knowledge
Incident management
Performance optimization

Tools

Kubernetes
GPU clusters
CUDA stack

Job description

Our client is a high-growth GPU infrastructure provider delivering factory-scale GPU-as-a-Service to organisations running large-scale training and high-throughput inference workloads.

They operate full-stack GPU environments end-to-end — including bare metal, Kubernetes control planes, storage, networking, observability, and day-two operations — enabling customers to deploy AI models with predictable performance and reliability.

The Role

This is a hands-on Technical Account Manager role supporting customers running production AI workloads on large-scale GPU clusters.

You will:

  • Act as the primary technical contact for a portfolio of customers
  • Troubleshoot production issues across GPU infrastructure and Kubernetes
  • Lead incident triage, mitigation, and communication
  • Support distributed training and inference workloads
  • Help customers optimise performance, reliability, and utilisation
  • Create documentation, runbooks, and best practices
  • Run regular technical check-ins and health reviews

This is an engineering-led TAM position — not a ticket-routing role.

Required Experience
  • 7+ years in customer-facing infrastructure roles
  • Strong hands-on Kubernetes experience (networking, storage, scaling, troubleshooting)
  • Experience operating bare metal GPU clusters
  • Linux systems and networking fundamentals
  • GPU drivers and CUDA stack familiarity
  • Experience supporting distributed training and/or inference workloads in production
  • Ability to lead incidents and communicate clearly under pressure
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GPU Infrastructure TAM — Production AI & Incident Lead
GPU Infrastructure TAM — Production AI & Incident Lead

asobbi • United States

Remote
USD 120,000 - 150,000
Senior Technical Account Manager
Senior Technical Account Manager

GMI Cloud • Mountain View (CA)

On-site
USD 140,000 - 180,000
Member of Technical Staff (Software Engineer, GPU Cluster Infrastructure)
Member of Technical Staff (Software Engineer, GPU Cluster Infrastructure)

United States Digital Space LLC • San Francisco (CA)

On-site
USD 180,000 - 260,000
Staff Technical Program Manager - AI Infrastructure
Staff Technical Program Manager - AI Infrastructure

Hamilton Barnes Associates Limited • New York (NY)

On-site
USD 200,000 - 240,000
Bonus
Equity
GPUaaS Kubernetes Platform Engineer
GPUaaS Kubernetes Platform Engineer

Veriipro • Irving (TX)

On-site
USD 140,000 - 180,000
Customer Solution Architect - Systems Integrator
Customer Solution Architect - Systems Integrator

Hamilton Barnes Associates Limited • New York (NY)

On-site
USD 225,000 - 275,000
RSU equity
20% bonus
AI Infra Engineer – SRE (Kubernetes)
AI Infra Engineer – SRE (Kubernetes)

Berrybytes • United States

On-site
USD 110,000 - 150,000
Member of Technical Staff (Software Engineer, GPU Cluster Infrastructure)
Member of Technical Staff (Software Engineer, GPU Cluster Infrastructure)

B Capital • United States

On-site
USD 180,000 - 230,000
Platform Engineer
Platform Engineer

Harrison Clarke • San Francisco (CA)

On-site
USD 120,000 - 160,000
AI Kernel / Cluster Engineer
AI Kernel / Cluster Engineer

Blue Signal Search • Santa Clara (CA)

On-site
USD 150,000 - 210,000