Principal Technical Account Manager, AWS Enterprise Support, NAMER-Sp

Amazon Web Services (AWS)

Chicago (IL)

On-site

USD 183,000 - 247,000

Full time

5 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Amazon Web Services, Inc. in NAMER region seeks a Principal AI/ML HPC Specialist to join the TAM team. You will lead strategic technical engagements, guide NAMER labs and enterprise customers through AI/ML transformation journeys, and architect scalable GPU/Trainium cluster solutions.

You'll work across Slurm, SageMaker HyperPod, DLAMIs, NCCL tuning, and cost governance, helping customers optimize training and inference at scale while partnering with TAMs and service teams.

Qualifications

  • 8+ years in Data & AI technologies including AI/ML, analytics, storage or databases.
  • 3+ years designing or consulting on large-scale ML training or inference architectures.
  • 10+ years IT development or consulting in software/cloud/AI/ML industries.
  • Ability to serve as a trusted technical advisor to enterprise customers.

Responsibilities

  • Lead deep technical engagements and optimize enterprise AI/ML workloads across GPU clusters.
  • Design production-grade AI/ML training and inference using Slurm, PyTorch FSDP, DeepSpeed, Megatron-LM, and SageMaker HyperPod.
  • Support customers in building ML solutions including LLMs, PINNs, and MLOps pipelines.
  • Improve training throughput, NCCL communications, and GPU utilization while ensuring governance and cost controls.
  • Partner with service teams to guide AI/ML infrastructure decisions spanning compute, networking, storage, and orchestration.

Skills

AI/ML Expertise
Cloud computing
Enterprise engagement

Tools

PyTorch
TensorFlow
MxNet
SageMaker HyperPod
ParallelCluster
DLAMI
NCCL
AWS Neuron SDK
Slurm

Job description

Description

As part of the AWS Applied AI Solutions organization, we have a vision to provide business applications, leveraging Amazon’s unique experience and expertise, that are used by millions of companies worldwide to manage day‑to‑day operations. We will accomplish this by accelerating our customers’ businesses through delivery of intuitive and differentiated technology solutions that solve enduring business challenges. We blend vision with curiosity and Amazon’s real‑world experience to build opinionated, turnkey solutions. Where customers prefer to buy over build, we become their trusted partner with solutions that are no‑brainers to buy and easy to use.

Are you ready to transform how businesses leverage artificial intelligence and machine learning at scale? Join our team and become a strategic partner in delivering Amazon AI/ML solutions that empower global enterprises to innovate, optimize, and achieve unprecedented operational excellence.

Amazon Web Services (AWS) is seeking an experienced Principal AI/ML HPC Specialist to join our Technical Account Manager (TAM) team. You’ll be at the forefront of solving complex AI HPC implementation challenges, guiding NAMER Resarch labs to enterprise customers through their most ambitious machine learning transformation journeys. By combining deep technical expertise with collaborative problem‑solving, you’ll help organizations unlock the full potential of artificial intelligence and machine learning technologies — from distributed model training on GPU clusters to production‑grade inference at scale.

AWS Support includes experts from across AWS who help our customers design, build, operate, and secure their cloud environments. Customers innovate with AWS Professional Services, upskill with AWS Training and Certification, optimize with AWS Support and Managed Services, and meet objectives with AWS Security Assurance Services. Our expertise and emerging technologies include AWS Partners, AWS Sovereign Cloud, AWS International Product, and AI/ML-native solutions. You’ll join a diverse team of technical experts in dozens of countries who help customers achieve more with the AWS cloud.

Key job responsibilities

  • Deliver Strategic Technical Engagements — Lead comprehensive technical deep‑dives and performance optimization for enterprise AI/ML workloads, including distributed training cluster architecture using AWS Parallel Computing Service (PCS) and AWS ParallelCluster, the latest GPU‑accelerated computing (i.e. P6/P6e , G7/G7e instances), AWS Trainium‑based training (Trn3 UltraServers), and multi‑node NCCL communication tuning over EFA’s SRD protocol.
  • Architect and Validate Innovative Solutions — Design and implement production‑grade AI/ML training and inference solutions leveraging Slurm‑based job scheduling, distributed training frameworks (PyTorch FSDP, DDP, DeepSpeed, Megatron‑LM), SageMaker HyperPod for managed GPU clusters with automated health checks and node replacement, high‑performance parallel storage (Amazon FSx for Lustre), and container runtimes on Deep Learning AMIs (DLAMIs) against reference architectures and HPC lens to ensure performance, reliability, and cost governance at scale. Architect solutions using P6e UltraServers for multi‑trillion parameter frontier models and Trn3 with the AWS Neuron SDK for cost‑optimized training and inference.
  • Enable Customer Success — Support customers in implementing business‑critical HPC capabilities, including the development of large language model (LLM) (Llama, GPT‑class models), physics‑informed neural networks (PINNs) and surrogate models, MLOps pipelines, simulation‑ML hybrid architectures orchestrated by AWS Step Functions and AWS Batch, distributed data processing, cluster observability, and governance controls for GPU/Trainium‑intensive workloads.
  • Enable Business Critical Outcomes — Partner with service teams to enhance model training throughput, optimize NCCL collective communications, improve GPU/Trainium utilization across multi‑node UltraClusters, and drive operational efficiency through proactive monitoring, automated failure recovery (HyperPod health checks), and capacity planning (EC2 Capacity Blocks for ML). Contribute to product roadmap PFR, share refrerence architecture, performance , and benchmarks with broader TAM and Technical communities.
  • Serve as Trusted Advisor and Advocate — Develop and nurture technical partnerships with enterprise stakeholders, serving as the trusted advisor for AI/ML infrastructure decisions spanning compute, networking (Elastic Fabric Adapter with SRD), storage, orchestration, and the HPC‑to‑AI convergence journey.

A day in the life

Your day will be dynamic and impactful, involving deep technical consultations on distributed training architectures, strategic solution design for GPU and Trainium cluster deployments, and collaborative problem‑solving across multi‑node ML environments. You’ll engage with technical leaders, architect innovative AI/ML implementations — from Slurm‑managed PCS clusters and SageMaker HyperPod to PyTorch FSDP/DeepSpeed training jobs and Neuron SDK compilation workflows — and provide expert guidance that bridges machine learning infrastructure with business objectives.

You will partner with TAMs, SAs, and service teams to provide customers with AWS AI/ML best practice guidance, diving deep into machine learning infrastructure services (PCS, ParallelCluster, HyperPod, Batch), promoting customers' AI/ML workloads to production, developing regional AI/ML strategies, advising on HPC‑to‑AI convergence patterns (simulation‑surrogate loops, physics‑informed neural networks), and training field teams on distributed training patterns, GPU/Trainium cluster operations, and the use cases and benefits of artificial intelligence and machine learning at scale.

About The Team

We are a collaborative group of technical innovators dedicated to pushing the boundaries of cloud computing and artificial intelligence. Our team thrives on solving complex challenges — from optimizing NCCL all‑reduce operations across hundreds of GPUs to architecting elastic training clusters that scale with customer demand. We believe in continuous learning, mutual support, and driving technological advancement.

Diverse Experiences

Amazon values diverse experiences. Even if you do not meet all of the preferred qualifications and skills listed in the job description, we encourage candidates to apply. If your career is just starting, hasn’t followed a traditional path, or includes alternative experiences, don’t let it stop you from applying.

Why AWS

Amazon Web Services (AWS) is the world’s most comprehensive and broadly adopted cloud platform. We pioneered cloud computing and never stopped innovating — that’s why customers from the most successful startups to Global 500 companies trust our robust suite of products and services to power their businesses.

Work/Life Balance

We value work‑life harmony. Achieving success at work should never come at the expense of sacrifices at home, which is why flexible work hours and arrangements are part of our culture. When we feel supported in the workplace and at home, there’s nothing we can’t achieve in the cloud.

Inclusive Team Culture

Here at AWS, it’s in our nature to learn and be curious. Our employee‑led affinity groups foster a culture of inclusion that empower us to be proud of our differences. Ongoing events and learning experiences, including our Conversations on Race and Ethnicity and AmazeCon conferences, inspire us to never stop embracing our uniqueness.

Mentorship and Career Growth

We’re continuously raising our performance bar as we strive to become Earth’s Best Employer. That’s why you’ll find endless knowledge‑sharing, mentorship and other career‑advancing resources here to help you develop into a better‑rounded professional.

Basic Qualifications
  • 8+ years of working with Data & AI related technologies, including, but not limited to, AI/ML, GenAI, Analytics, Database, and/or Storage experience
  • 3+ years of hands‑on experience designing, implementing, or consulting on large‑scale ML training or inference architectures in a customer‑facing role
  • 10+ years of IT development or implementation/consulting in the software, cloud computing, or AI/ML industries
  • Demonstrated ability to serve as a trusted technical advisor to enterprise customers
Preferred Qualifications
  • Experience with deep learning libraries such as PyTorch, TensorFlow, MxNet Research publications in computer vision, deep learning or machine learning at peer‑reviewed workshops, conferences or journals
  • Experience with training and deploying machine learning systems to solve large‑scale optimizations
  • 5+ years of solving problems with technology in the Healthcare/Life Sciences Industry experience
  • Experience in the financial services industry

Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.

Los County applicants: Job duties for this position include: work safely and cooperatively with other employees, supervisors, and staff; adhere to standards of excellence despite stressful conditions; communicate effectively and respectfully with employees, supervisors, and staff to ensure exceptional customer service; and follow all federal, state, and local laws and Company policies. Criminal history may have a direct, adverse, and negative relationship with some of the material job duties of this position. These include the duties and responsibilities listed above, as well as the abilities to adhere to company policies, exercise sound judgment, effectively manage stress and work safely and respectfully with others, exhibit trustworthiness and professionalism, and safeguard business operations and the Company’s reputation. Pursuant to the Los Angeles County Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records.

Pursuant to the San Francisco Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records.

Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit https://amazon.jobs/content/en/how-we-hire/accommodations for more information. If the country/region you’re applying in isn’t listed, please contact your Recruiting Partner.

The base salary range for this position is listed below. Your Amazon package will include sign‑on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at https://amazon.jobs/en/benefits.

USA, CA, San Francisco - 210,200.00 - 284,300.00 USD annually

USA, CA, Santa Clara - 210,200.00 - 284,300.00 USD annually

USA, IL, Chicago - 182,800.00 - 247,300.00 USD annually

USA, NY, New York - 201,000.00 - 272,000.00 USD annually

USA, TX, Austin - 182,800.00 - 247,300.00 USD annually

USA, TX, Dallas - 182,800.00 - 247,300.00 USD annually

USA, WA, Seattle - 182,800.00 - 247,300.00 USD annually

Company

Amazon Web Services, Inc.

Job ID: A10535213

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Principal Technical Account Manager, AWS Enterprise Support, NAMER-Sp
Principal Technical Account Manager, AWS Enterprise Support, NAMER-Sp

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 183,000 - 247,000
RSUs
Health insurance
Principal Technical Account Manager, AWS Enterprise Support, NAMER-Sp
Principal Technical Account Manager, AWS Enterprise Support, NAMER-Sp

Amazon Web Services (AWS) • Santa Clara (CA)

On-site
USD 210,000 - 284,000
Health insurance
401(k) matching
Paid time off
+1
Principal Technical Account Manager, AWS Enterprise Support, NAMER-Sp
Principal Technical Account Manager, AWS Enterprise Support, NAMER-Sp

Amazon Web Services (AWS) • New York (NY)

On-site
USD 201,000 - 272,000
Health insurance
RSUs
401(k) matching
+2
Principal Technical Account Manager, AWS Enterprise Support, NAMER-Sp
Principal Technical Account Manager, AWS Enterprise Support, NAMER-Sp

Amazon Web Services (AWS) • Austin (TX)

On-site
USD 183,000 - 247,000
Principal Technical Account Manager, AWS Enterprise Support, NAMER-Sp
Principal Technical Account Manager, AWS Enterprise Support, NAMER-Sp

Amazon • New York (NY)

On-site
USD 201,000 - 272,000
Principal Technical Account Manager, AWS Enterprise Support, NAMER-Sp
Principal Technical Account Manager, AWS Enterprise Support, NAMER-Sp

Amazon • Chicago (IL)

On-site
USD 183,000 - 247,000
Principal AI/ML HPC Specialist Technical Account Manager (STAM) , AWS Enterprise Support, NAMER-Sp
Principal AI/ML HPC Specialist Technical Account Manager (STAM) , AWS Enterprise Support, NAMER-Sp

Amazon Web Services (AWS) • Herndon (VA)

On-site
USD 183,000 - 247,000
Principal AI/ML HPC Specialist Technical Account Manager (STAM) , AWS Enterprise Support, NAMER-Sp
Principal AI/ML HPC Specialist Technical Account Manager (STAM) , AWS Enterprise Support, NAMER-Sp

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 183,000 - 247,000
Principal AI/ML HPC Specialist Technical Account Manager (STAM) , AWS Enterprise Support, NAMER-Sp
Principal AI/ML HPC Specialist Technical Account Manager (STAM) , AWS Enterprise Support, NAMER-Sp

Amazon • Seattle (WA)

On-site
USD 183,000 - 247,000
Sr. Delivery Consultant - AI/ML, AWS Professional Services
Sr. Delivery Consultant - AI/ML, AWS Professional Services

Amazon Web Services (AWS) • Arlington (VA)

On-site
USD 154,000 - 208,000