System Design Engineer - AI Cluster Software Engineer

AMD

Austin (TX)

On-site

USD 140,000 - 200,000

Full time

12 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

AMD in Austin, TX is seeking a hands-on full-stack developer to create tools for design and deployment of large-scale AI/ML clustered infrastructure. You will work with cutting-edge agentic tools to develop, deploy, and maintain these applications, joining a growing team of engineers across domains.

The role requires strong Linux fundamentals, experience with frontend and backend development, CI/CD, and securing applications, with collaboration across stakeholders in an Agile environment.

Qualifications

  • Professional software development experience, including substantial experience building and supporting web applications.
  • Demonstrated use of AI coding assistants and LLM-powered developer tools.
  • Experience deploying and operating applications using Linux, containers, CI/CD, and cloud or on-premises infrastructure.
  • Experience with automated testing, source control, code review, debugging, and production support.
  • Working knowledge of web application security, authentication, authorization, and secure secrets handling.

Responsibilities

  • Translate requirements into flexible, future-proof design solutions.
  • Develop, iterate, and maintain tools and applications for large-scale AI cluster design and deployment.
  • Interface code between third-party tools.
  • Own features from requirements through deployment and maintenance.
  • Collaborate in Agile sprints with stakeholders across domains.
  • Participate in code reviews and retros.

Skills

AI tools & LLMs
Frontend development (JavaScript/TypeS
Backend services & APIs
Linux & Containers
CI/CD
Code reviews & testing
Security & secrets handling

Education

Bachelor's or Master's in CS or Software/CS Engineering

Tools

React/Angular/Vue
Kubernetes
Slurm
Apptainer/Singularity
Terraform/Ansible
Python
Bash

Job description

WHAT YOU DO AT AMD CHANGES EVERYTHING

At AMD, our mission is to build great products that accelerate next-generation computing experiences—from AI and data centers, to PCs, gaming and embedded systems. Grounded in a culture of innovation and collaboration, we believe real progress comes from bold ideas, human ingenuity and a shared passion to create something extraordinary. When you join AMD, you’ll discover the real differentiator is our culture. We push the limits of innovation to solve the world’s most important challenges—striving for execution excellence, while being direct, humble, collaborative, and inclusive of diverse perspectives. Join us as we shape the future of AI and beyond.

Together, we advance your career.
THE ROLE:

This is a hands-on role for a full-stack developer to create and deliver a range of tools and applications focused on design and deployment of large-scale AI/ML clustered infrastructure. You will be working with the latest agentic tools and patterns to develop, deploy, and maintain these applications. You’ll join a growing team of multi-disciplined engineers that operates across industry verticals as subject matter experts in the AI stack and across the cluster.

THE PERSON:
  • Demonstrated use of AI coding assistants and LLM-powered developer tools: daily user of AI agents and tools
  • Professional software development experience, including substantial experience building and supporting web applications
  • Proficiency in modern frontend development using JavaScript or TypeScript and a framework such as React, Angular, or Vue
  • Experience developing backend services and APIs using a modern server-side language or framework
  • Experience deploying and operating applications using Linux, containers, CI/CD, and cloud or on-premises infrastructure
  • Experience with automated testing, source control, code review, debugging, and production support
  • Working knowledge of web application security, authentication, authorization, and secure secrets handling
KEY RESPONSIBILITIES:
  • Partner with engineering peers, domain experts in adjacent teams, and business stakeholders to understand requirements and translate them into flexible, future-proof design solutions
  • Hands on development, iteration, and maintenance of tools and applications that codify various aspects of large-scale AI cluster design stages and cluster deployment activities
  • Design and development of cohesive interface code between disparate third party tools
  • Own features from requirements and design through deployment and ongoing maintenance
  • Work in an iterative software environment, including planning and delivering work in small increments, collaborating with stakeholders, often in different areas of domain expertise (Agile development practices)
  • Participate in code reviews and retros; adapt to changing requirements and priorities
PREFERRED EXPERIENCE:
  • Strong Linux fundamentals: Linux operating systems, networking, filesystems, containers, performance tooling (perf, flamegraphs, nvprof/rocprof, basic eBPF).
  • Clear communication: ability to turn complex systems into accessible, structured documentation with diagrams and reproducible steps
  • AMD ecosystem experience: ROCm, RCCL, Instinct GPUs, EPYC platforms, compiler/toolchain impacts, and performance tuning
  • Orchestration models: Slurm configuration patterns, Kubernetes for HPC/AI (GPU operators, device plugins), Apptainer/Singularity
  • Automation, IaC , and scripting tools/languages (Ansible, Terraform, Python, bash)
  • Storage/data: knowledge of or familiarity with parallel filesystems (Lustre, BeeGFS), object stores, RDMA, data pipeline throughput and caching strategies
  • Hands-on familiarity with on-premises infrastructure, particularly for AI/ML/HPC workloads would be beneficial
ACADEMIC CREDENTIALS:
  • Bachelors or Masters degree in computer science or software/computer engineering
LOCATIONS:
  • Austin, TX
  • Santa Clara, CA
  • Secaucus, NJ
  • Seattle, WA

This role is not eligible for visa sponsorship.

Benefits offered are described: AMD benefits at a glance.

AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process.

AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here.

This posting is for an existing vacancy.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

System Design Engineer - AI Cluster Software Engineer
System Design Engineer - AI Cluster Software Engineer

Advanced Micro Devices • Austin (TX)

Hybrid
USD 150,000 - 230,000
AMD benefits
Hybrid work model
System Design Engineer - AI Cluster Software Engineer
System Design Engineer - AI Cluster Software Engineer

AMD • Santa Clara (CA)

On-site
USD 170,000 - 250,000
AMD Benefits at a glance
System Design Engineer - AI Cluster Software Engineer
System Design Engineer - AI Cluster Software Engineer

Advanced Micro Devices • Santa Clara (CA)

On-site
USD 140,000 - 230,000
Gen AI Software Development Engineer
Gen AI Software Development Engineer

Advanced Micro Devices • Santa Clara (CA)

On-site
USD 140,000 - 210,000
Sr. AI/ML Engineer
Sr. AI/ML Engineer

Advanced Micro Devices • Austin (TX)

Hybrid
USD 140,000 - 210,000
GenAI Software Development Engineer
GenAI Software Development Engineer

Advanced Micro Devices • Santa Clara (CA), Northern (KY)

Hybrid
USD 180,000 - 260,000
Gen AI Software Development Engineer
Gen AI Software Development Engineer

Jobzhr • Santa Clara (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Sr. AI/ML Engineer
Sr. AI/ML Engineer

AMD • Austin (TX)

Hybrid
USD 140,000 - 200,000
AMD benefits
AI Systems Engineer - HPC
AI Systems Engineer - HPC

Advanced Micro Devices • San Jose (CA), Northern (KY)

Hybrid
USD 140,000 - 190,000
Forward Deployed AI Engineer
Forward Deployed AI Engineer

Advanced Micro Devices • Santa Clara (CA)

Hybrid
USD 180,000 - 240,000
AMD Benefits