Principal Software Engineer — AI Performance & Reliability

Advanced Micro Devices, Inc.

San Jose (CA)

Hybrid

USD 250,000 - 420,000

Full time

48 hours ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Hybrid work model
Benefits package

Job summary

Advanced Micro Devices, Inc. is seeking a Principal or Fellow level software engineer to advance AI infrastructure, focusing on performance, reliability and scalable workloads for model training and inference.

You will collaborate with customers and cross‑functional teams to identify bottlenecks, implement reusable solutions, and shape product capabilities to run demanding AI workloads efficiently at scale.

Qualifications

  • PhD or master’s degree in AI, ML, CS or related field or equivalent experience.
  • Proven track record in optimizing AI workloads and production systems.
  • Experience with large-scale model training and inference pipelines.

Responsibilities

  • Profile and optimize AI model training and inference workloads.
  • Improve throughput, latency, memory efficiency and scalability.
  • Identify bottlenecks across models, frameworks, runtimes and hardware.
  • Develop performance tools, benchmarks, automation and observability.
  • Collaborate with customers and internal teams to resolve production issues.

Skills

AI model training
Model inference
Performance profiling
Python
C++
PyTorch
TensorFlow
CUDA

Education

PhD or Masters

Tools

PyTorch
TensorFlow
JAX
CUDA
XLA
MLIR
NCCL

Job description

ADVANCE YOUR CAREER. ADVANCE THE WORLD.

At AMD, we believetechnology has the power to solve the world’s most important challenges. From advancing healthcare and scientific discovery to powering AI and the technologies people rely on every day, innovation at AMDis shapingthefuture.

Whetheryou’redesigning next-gen processors, enabling AI breakthroughs, orbringing leading edge products to market, every role at AMD contributes to something bigger— technologythat moves the world forward.Join us and, together, we’ll advance your career.

THE ROLE:

We are looking for a strong, Principal or Fellow level software engineer to join our AI Infrastructure team. You will work on improving the performance, efficiency, and reliability of AI workloads across both model training and inference.

Our team supports a broad range of machine learning systems, including large language models, diffusion models, and recommendation models. You will collaborate closely with customers and internal engineering teams to understand performance bottlenecks, optimize workloads, and ensure that models run reliably at scale.

This role is a strong fit for an engineer who enjoys working across the AI software and hardware stack, solving technically challenging performance problems, and partnering directly with customers to make them successful.

You will help customers achieve meaningful improvements in model performance and system reliability. You will identify difficult bottlenecks, develop reusable solutions, and help shape the infrastructure and product capabilities needed to run demanding AI workloads efficiently at scale.

THE PERSON:
  • Profile and optimize AI model training and inference workloads.
  • Improve model throughput, latency, memory efficiency, scalability, and reliability.
  • Identify bottlenecks across models, frameworks, compilers, runtimes, operating systems, and hardware.
  • Optimize workloads involving large language models, diffusion models, recommendation systems, and other modern machine learning architectures.
  • Develop performance tooling, benchmarks, automation, and observability systems.
  • Investigate and resolve complex production issues affecting AI workloads.
  • Collaborate with customers to understand their technical requirements, reproduce issues, and recommend effective solutions.
  • Translate customer feedback into product and infrastructure improvements.
  • Work closely with machine learning engineers, systems engineers, hardware teams, and product teams.
  • Document performance findings, technical recommendations, and best practices.
KEY RESPONSIBILITIES:
  • Strong software engineering skills and experience building production-quality systems.
  • Experience working with AI infrastructure for model training, inference, or both.
  • Demonstrated experience profiling and optimizing machine learning models or AI workloads.
  • Strong foundations in computer architecture, including processors, memory hierarchies, parallelism, and performance tradeoffs.
  • Solid understanding of systems performance concepts such as latency, throughput, memory bandwidth, utilization, and distributed communication.
  • Proficiency in languages such as Python, C++, or similar systems-oriented programming languages.
  • Experience with machine learning frameworks such as PyTorch, TensorFlow, or JAX.
  • Strong debugging and analytical skills, with the ability to investigate problems across multiple layers of the technology stack.
  • Clear written and verbal communication skills.
  • A customer-focused mindset and willingness to work directly with customers through technical evaluations, deployments, troubleshooting, and ongoing support.
PREFERRED EXPERIENCE:
  • Experience optimizing large language models, diffusion models, or recommendation models.
  • Experience with GPU, accelerator, or distributed computing environments.
  • Familiarity with technologies such as ROCm, HIP, CUDA, Triton, XLA, MLIR, NCCL, or similar performance-oriented tools and runtimes.
  • Experience with distributed training, model serving, quantization, compilation, kernel optimization, or memory optimization.
  • Experience operating AI systems in production environments.
  • Prior experience in solutions engineering, field engineering, developer relations, or another customer-facing technical role.
  • Experience designing benchmarks and conducting systematic performance analysis.
ACADEMIC CREDENTIALS:
  • A PhD (or a master’s degree with equivalent experience) in artificial intelligence, machine learning, computer science, or a related field.

LOCATION:

San Jose, CA or Bellevue, WA preferred (Hybrid). Other US locations may be considered.

#LI-MV1

#HYBRID

  • Benefits offered are described: AMD benefits at a glance.

AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process.

AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here.

This posting is for an existing vacancy.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Fellow Software Engineer - AI Performance & Reliability
Fellow Software Engineer - AI Performance & Reliability

Advanced Micro Devices • San Jose (CA)

On-site
USD 180,000 - 240,000
Fellow Software Engineer — AI Performance & Reliability
Fellow Software Engineer — AI Performance & Reliability

AMD • San Jose (CA)

On-site
USD 180,000 - 260,000
Frontier AI Workloads - Performance and Scalability Engineer
Frontier AI Workloads - Performance and Scalability Engineer

AMD • San Jose (CA)

On-site
USD 150,000 - 200,000
Competitive salary
Health benefits
Career advancement opportunities
Principal Engineer, Efficient GenAI
Principal Engineer, Efficient GenAI

Advanced Micro Devices, Inc. • San Jose (CA)

On-site
USD 180,000 - 250,000
AI Research Scientist
AI Research Scientist

Advanced Micro Devices, Inc. • Bellevue (WA)

On-site
USD 180,000 - 280,000
Benefits at a glance
Principal AI Performance and Tools
Principal AI Performance and Tools

AMD • San Jose (CA)

Hybrid
USD 150,000 - 190,000
AMD benefits at a glance
Principal AI Performance and Tools
Principal AI Performance and Tools

Advanced Micro Devices, Inc. • San Jose (CA), Northern (KY)

Hybrid
USD 180,000 - 260,000
AMD Benefits at a glance
Principal Engineer, Efficient GenAI
Principal Engineer, Efficient GenAI

AMD • San Jose (CA)

On-site
USD 250,000 - 360,000
Fellow GPU Performance Engineer AI Training at Scale
Fellow GPU Performance Engineer AI Training at Scale

Advanced Micro Devices, Inc. • San Jose (CA)

On-site
Confidential
Fellow GPU Performance Optimization Engineer
Fellow GPU Performance Optimization Engineer

Advanced Micro Devices, Inc. • San Jose (CA)

On-site
USD 150,000 - 180,000