Principal Software Developer – AI/ML Performance Validation & Systems Testing

AMD

San Jose (CA)

Hybrid

USD 190,000 - 240,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Benefits at a glance

Job summary

AMD in San Jose seeks a Principal Software Engineer to lead ROCm validation across compute workloads and server-class systems. You will define testing strategy, build scalable test infrastructure, and mentor teams to ensure high-quality ROCm releases.

Join a leadership role focused on multi-node validation, LLM training workloads, and performance benchmarking for next-generation Instinct GPUs, in a hybrid San Jose environment.

Qualifications

  • Experience leading complex systems validation and quality engineering.
  • Expert Python for test automation and infrastructure; strong C++ for debugging and extending production code.
  • Deep validation expertise in GPU software stacks, AI/ML frameworks, HPC runtimes, and Linux kernel/drivers.

Responsibilities

  • Own end-to-end ROCm validation architecture across multi-GPU generations and server platforms.
  • Define release-qualification gates and exit criteria for ROCm software releases.
  • Architect the test infrastructure, including distributed test runners, CI fleets, and data lakes.

Skills

Python
C++
Validation
Linux internals
GPU software stacks
CI/CD
HPC/Distributed systems

Education

BS in Computer Science or Computer Engineering

Tools

rocprof
Nsight
Omniperf
GitHub Actions
Jenkins

Job description

ADVANCE YOUR CAREER. ADVANCE THE WORLD.

At AMD, we believetechnology has the power to solve the world’s most important challenges. From advancing healthcare and scientific discovery to powering AI and the technologies people rely on every day, innovation at AMDis shapingthefuture.

Whetheryou’redesigning next-gen processors, enabling AI breakthroughs, orbringing leading edge products to market, every role at AMD contributes to something bigger— technologythat moves the world forward.Join us and, together, we’ll advance your career.

THE ROLE:

Weare seeking aPrincipal Software Engineertoserve as the senior technical leader forROCm software validationacrosscompute workloads and server-class systems. Inthis individual-contributor leadership role, you will define howAMD provesROCm is ready to ship— from unit andcomponenttesting, through full-stack workload validation, to multi-node system-level qualification on AMD Instinct™ GPU platforms. Youwill set the technical direction for validation strategy, build and evolve the test infrastructure thatgates everyROCm release, and personally drive the hardestdebugging, characterization, and qualification problems. Your work directly determines thequality bar experienced by hyperscalers, OEMs, sovereign-AI customers, and the open-source community runningROCm inproduction.

THE PERSON:

Youwill set the technical direction for validation strategy, build and evolve the test infrastructure thatgates everyROCm release, and personally drive the hardestdebugging, characterization, and qualification problems. Your work directly determines thequality bar experienced by hyperscalers, OEMs, sovereign-AI customers, and the open-source community runningROCm inproduction.

KEY RESPONSIBILITIES:
  • Ownthe end-to-end validation architecture for ROCm — unit, integration, framework, workload, performance, stress, stability, scale-out, and system-level test layers — across multiple GPU generations and server platforms.
  • Define release-qualification gates and exit criteria for ROCm software releases (functional coverage, performance regressions, stability hours, scale targets, RAS criteria) and drive the org to meet them.
  • Architect the test infrastructure — distributed test runners, GitHub Actions/ Jenkins / internal CI fleets, hardware lab orchestration, resultdatalakes, flaky-test detection, bisectionautomation, and self-servicedeveloper pre-submit pipelines.
  • Champion modern, agile quality engineering — shift-left testing, test pyramids, contract testing betweenlayers, hermetic test environments, deterministic reproducers, and continuous validation intrunk.
  • Set the bar for GitHub-based quality workflows — PR gatingpolicy, requiredchecks, code-coverage standards, bug-bashandtriage cadences, and disciplined issue management acrossROCm/*repositories and partner upstream projects.
  • Lead complex escalation debug — partner with development, hardware, firmware, and customer-facing teams to root-cause the hardest multi-day, multi-node, multi-component failures and convert findings into durable test coverage.
  • Influence the roadmap— work with product management, silicon, platform, and softwarearchitecture to ensure validation readiness fornext-generation Instinct GPUs and serverplatformsbefore tape-inmilestones and silicon arrival.
  • Mentor and elevate Senior and Staff validation engineers, SDETs, and SQA leads; raise the technical bar through designreview, code review, and written guidance.
  • Represent ROCm validation externally — strategic customer engagements, OEM qualification programs, and open-source community quality initiatives.
  • Lead system-level testing for server nodes— multi-GPU topologies, PCIe/InfinityFabric/xGMI, BMC/IPMI, thermal/power, firmware interactions, and multi-node fabric(Ethernet/InfiniBand/UALink) bring-up andvalidation.Drive compute workload validation and characterization— LLM training andinference(PyTorch, vLLM, Triton, JAX), recommender systems, scientific HPC kernels, MLPerf-class benchmarks— establishing reproducible methodology, baselines, and regression tracking.
PREFERRED EXPERIENCE:
  • Software engineering experience in validation, SDET, or quality engineering, including experience leading complex systems validation.
  • Expert Python for test automation and infrastructure; strong C++ for debugging and extending production code.
  • Deep validation expertise in two or more of the following:
    • GPU software stacks (ROCm, CUDA, oneAPI, SYCL)
    • AI/ML frameworks (PyTorch, TensorFlow, JAX, Triton, vLLM)
    • HPC runtimes and communication libraries (MPI, RCCL/NCCL, UCX, Libfabric)
    • Linux kernel, GPU drivers, or accelerator firmware
    • Distributed systems and large-scale cluster software
  • Experience validating multi-GPU, multi-node server platforms, including stress, soak, fault injection, and RAS testing.
  • Experience defining and delivering release qualification programs for hyperscalers, OEMs, or Tier-1 customers.
  • Contributions to validation, CI, or test infrastructure for ROCm, PyTorch, LLVM, Triton, vLLM, or similar open-source projects.
  • Experience leading adoption of agentic AI workflows, including automated testing, AI-driven debugging, MCP, and RAG-based engineering solutions.
  • Experience validating or operating large-scale GPU clusters (256+ GPUs), including fabric bring-up, health monitoring, and diagnostics.
  • Familiarity with AI training, inference, and HPC benchmark methodologies.
  • Experience with performance validation, profiling tools (rocprof, Omniperf, Nsight), and regression analysis.
  • Familiarity with hardware lab automation, including BMC/IPMI/Redfish, PDU control, serial consoles, automated re-imaging, and topology-aware scheduling.
  • Experience supporting validation for pre-silicon, emulation, and first-silicon accelerator bring-up.
ACADEMIC CREDENTIALS:
  • BS/MS/PhDin Computer Science, Computer Engineering, orrelated discipline (or equivalent demonstrated experience).
LOCATION: San Jose, California

#LI-DR1

#LI-HYBRID

Benefits offered are described: AMD benefits at a glance.

AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process.

AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here.

This posting is for an existing vacancy.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Principal Software Developer – AI/ML Performance Validation & Systems Testing
Principal Software Developer – AI/ML Performance Validation & Systems Testing

Advanced Micro Devices • San Jose (CA), Northern (KY)

Hybrid
USD 200,000 - 260,000
Senior AI Platform Development & Validation Engineer
Senior AI Platform Development & Validation Engineer

Advanced Micro Devices • Secaucus (NJ)

On-site
USD 150,000 - 210,000
AI Systems Validation Engineer
AI Systems Validation Engineer

AMD • Secaucus (NJ)

On-site
USD 150,000 - 210,000
Princiapl AI Validation and Test Engineer
Princiapl AI Validation and Test Engineer

AMD • Secaucus (NJ)

On-site
USD 140,000 - 200,000
AMD benefits at a glance
AI Systems Validation Engineer
AI Systems Validation Engineer

Advanced Micro Devices • Secaucus (NJ)

On-site
USD 140,000 - 230,000
Principal System Test Validation Engineer
Principal System Test Validation Engineer

Advanced Micro Devices • Austin (TX)

On-site
USD 140,000 - 180,000
AI Validation and Test Engineer
AI Validation and Test Engineer

AMD • Secaucus (NJ)

On-site
USD 150,000 - 210,000
Benefits package
AI Validation and Test Engineer
AI Validation and Test Engineer

Advanced Micro Devices • Secaucus (NJ)

On-site
USD 150,000 - 210,000
AMD benefits
Lead AI Validation and Test Engineer
Lead AI Validation and Test Engineer

Advanced Micro Devices • Austin (TX)

On-site
USD 180,000 - 240,000
Principal System Test Validation Engineer
Principal System Test Validation Engineer

AMD • Austin (TX)

On-site
USD 150,000 - 190,000
AMD Benefits