Cloud & Customer Solutions Engineer - DC GPU

Advanced Micro Devices

Bellevue, Northern (WA, KY)

Hybrid

USD 150,000 - 190,000

Full time

46 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Advanced Micro Devices (AMD) is seeking a Cloud and Customer Solutions Engineer for its Applied AI team. You will embed with AI customers to take AMD Instinct GPU clusters from delivery to sustained production, owning outcomes end-to-end, and translating field learnings into durable upstream improvements.

You will debug complex distributed systems, write production code, and work across ROCm, serving frameworks, and cloud/bare-metal deployments to ensure reliable performance for enterprise

Qualifications

  • 5+ years of production software or infrastructure engineering, including field deployments
  • Hands-on experience with GPU compute at scale and production workloads
  • Strong knowledge of modern AI infrastructure stacks (Kubernetes/Slurm, containerized GPU workloads)
  • Cloud platform experience (AWS/Azure/GCP) and hybrid/bare-metal deployments
  • Proficiency in Python and a systems language; ability to modify large codebases

Responsibilities

  • Own customer deployments end-to-end, from bring-up to production readiness
  • Deploy and tune large-scale training/inference stacks against customer workloads across cloud and bare-metal environments
  • Lead root-cause analysis and resolve production incidents
  • Provide field learnings to improve ROCm, serving frameworks, and reference architectures
  • Transfer operational capability to customer teams with runbooks and enablement
  • Contribute upstream to ROCm issues and patches and provide structured feedback to product teams

Skills

GPU compute at scale
Kubernetes
Slurm
RCCL/NCCL
Observability tooling
Python
Systems language
Cloud platforms
On-site customer engagements
CUDA ecosystem experience
Open-source contributions

Education

Bachelor's or Master's degree in CS/CE/EE

Tools

Prometheus
Grafana
ROCm
CUDA ecosystem

Job description

ADVANCE YOUR CAREER. ADVANCE THE WORLD.

At AMD, we believetechnology has the power to solve the world’s most important challenges. From advancing healthcare and scientific discovery to powering AI and the technologies people rely on every day, innovation at AMDis shapingthefuture.

Whetheryou’redesigning next-gen processors, enabling AI breakthroughs, orbringing leading edge products to market, every role at AMD contributes to something bigger— technologythat moves the world forward.Join us and, together, we’ll advance your career.

THE TEAM:

AMD's Data Center GPU organization is transforming the industry with our AI based Graphic Processors. Our primary objective is to design exceptional products that drive the evolution of computing experiences, serving as the cornerstone for enterprise Data Centers, (AI) Artificial Intelligence, HPC and Embedded systems. If this resonates with you, come and joining our Data Center GPU organization where we are building amazing AI powered products with amazing people.

THE ROLE:

As a Cloud and Customer Solutions Engineer on AMD's Applied AI team, you will embed directly with AMD's most strategic AI customers — frontier labs, NeoCloud providers, CSPs, and AI-native companies — to take AMD Instinct GPU clusters from delivery to sustained production excellence. You own the customer outcome end-to-end: cluster bring-up and certification, workload deployment and performance, production incident response, and the transfer of operational capability that moves customers toward autonomous operation of their AMD fleets.

To be direct about what this role is: despite the "Solutions" title, this is not a pre-sales or demo role. You will write production code, operate live clusters, carry accountability for customer production outcomes, and be the engineer in the room when things break at scale. What you learn in the field, you convert into durable improvements — to ROCm, to the open-source serving ecosystem, and to the reference architectures every subsequent deployment inherits.

THE PERSON:

You are a strong production engineer who is energized rather than drained by ambiguity, customer pressure, and environments you do not control. You can debug a distributed training hang at 2am, explain the root cause to a customer VP at 9am, and land the fix upstream by the end of the week. You measure success by customer production outcomes, not code merged or tickets closed. When something is broken on a cluster you touch, it is your problem until it is fixed or explicitly handed off.

KEY RESPONSIBILITIES:

  • Own customer deployments end-to-end: cluster bring-up and burn-in, production readiness certification, workload onboarding, performance validation, and sustained production operation on AMD Instinct GPU fleets
  • Deploy and tune large-scale training and inference stacks (ROCm, vLLM, SGLang, RCCL, Kubernetes, Slurm) against customer‑specific workloads and SLOs across cloud, NeoCloud, and bare‑metal environments
  • Lead root‑cause analysis and resolution of production incidents on customer clusters, including Sev‑1 response, and drive fixes to permanent closure
  • Deploy agentic AI solutions into customer environments in partnership with Agentic Data Engineers, and own their production behavior within the engagement
  • Build the observability, benchmarking, and validation tooling needed to certify clusters as production‑ready and keep them there
  • Transfer operational capability to customer teams — documentation, runbooks, and hands‑on enablement — moving customers up the operator‑autonomy ladder from assisted operation to independent production ownership
  • Contribute field learnings to the Applied AI team's skills library and engagement memory databases, so deployment knowledge compounds across the practice
  • Convert field findings into upstream contributions — ROCm issues and patches, serving‑framework improvements, reference‑architecture updates — and provide structured field signal to AMD product, software, and silicon teams

PREFERRED EXPERIENCE:

  • 5+ years of production software or infrastructure engineering, including significant time operating or deploying systems in environments you did not build (level flexible for exceptional candidates)
  • Hands‑on experience with GPU compute at scale: cluster deployment, distributed training or high‑throughput inference, performance debugging, and workload optimization
  • Strong working knowledge of the modern AI infrastructure stack: Kubernetes and/or Slurm, containerized GPU workloads, collective communication libraries (RCCL/NCCL), high‑performance networking (RoCE/InfiniBand), and observability tooling (Prometheus, Grafana)
  • Cloud platform depth (AWS, Azure, GCP, or NeoCloud environments), including hybrid and bare‑metal deployment patterns
  • Proficiency in Python and at least one systems language; comfort navigating and modifying large codebases you did not write
  • Working familiarity with LLM application patterns — inference serving, RAG, and agentic workflows — sufficient to deploy and troubleshoot them in customer environments
  • Direct customer‑facing experience: embedded deployments, technical escalations, on‑site engagements, or equivalent
  • Experience with ROCm and AMD Instinct GPUs strongly preferred; deep CUDA‑ecosystem experience with demonstrated ability to work cross‑platform also valued
  • Open‑source contribution history in AI/ML infrastructure projects is a plus

PREFERRED ACADEMIC CREDENTIALS:

  • Bachelor's or Master's degree in Computer Science, Computer Engineering, Electrical Engineering, or equivalent practical experience

This role is not eligible for visa sponsorship.

#LI-RW1

#LI-HYBRID

Benefits offered are described: AMD benefits at a glance.

AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee‑based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third‑party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process.

AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here.

This posting is for an existing vacancy.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI & Datacenter Solutions Architect
AI & Datacenter Solutions Architect

AMD • South Carolina

On-site
USD 180,000 - 260,000
Senior Field Application Engineer – AI
Senior Field Application Engineer – AI

Socket.dev • Austin (TX)

Hybrid
USD 140,000 - 190,000
Benefits at a glance
System Design Engineer - AI Cluster Software Engineer
System Design Engineer - AI Cluster Software Engineer

Advanced Micro Devices • Santa Clara (CA)

On-site
USD 140,000 - 230,000
Systems Application Engineer
Systems Application Engineer

AMD • Austin (TX)

Hybrid
USD 120,000 - 180,000
System Design Engineer - AI Cluster Software Engineer
System Design Engineer - AI Cluster Software Engineer

AMD • Santa Clara (CA)

On-site
USD 170,000 - 250,000
AMD Benefits at a glance
Director, Program Management - DC GPU
Director, Program Management - DC GPU

Advanced Micro Devices • Austin (TX), Northern (KY)

Hybrid
USD 190,000 - 230,000
Sr. Manager Agentic AI / Data Software Development - DC GPU
Sr. Manager Agentic AI / Data Software Development - DC GPU

Advanced Micro Devices • San Jose (CA)

Hybrid
USD 140,000 - 190,000
AI Cluster Technical Program Manager – Validation, Debug & Agentic AI
AI Cluster Technical Program Manager – Validation, Debug & Agentic AI

AMD • United States

On-site
USD 140,000 - 180,000
Data Center Infrastructure Architect
Data Center Infrastructure Architect

CareerArc • North Carolina

On-site
USD 150,000 - 230,000
Corp. VP Customer Engineering DCGPU
Corp. VP Customer Engineering DCGPU

Advanced Micro Devices • Austin (TX)

On-site
USD 300,000 - 600,000