Summer 2027 PhD AI Research Infrastructure, RL Post-Training Intern

Advanced Micro Devices, Inc.

Santa Clara (CA)

Hybrid

USD 60,000 - 100,000

Full time

4 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Advanced Micro Devices, Inc. (AMD) is seeking an AI Research Intern in Santa Clara, CA to contribute to RL-based training infrastructure for large language and multimodal models.

You will work with researchers to design scalable systems, optimize performance, and build tools for experiments, logging, and reproducibility across distributed GPU clusters.

Qualifications

  • Pursuing a PhD in CS/ML/AI or related field.
  • Strong programming in Python and PyTorch.
  • Knowledge of RL, LLM post-training, RLHF/RLAIF, or preference optimization.
  • Experience with distributed training and multi-GPU workloads.
  • Familiarity with data, tensor, pipeline or sequence parallelism.
  • Experience with training/inference frameworks and cluster environments.
  • Understanding GPU performance, memory, networking, and distributed communication.
  • Experience building research infrastructure, profiling, or debugging distributed workloads.
  • Familiarity with containerization and experiment tracking is a plus.

Responsibilities

  • Develop and optimize infrastructure for RL-based post-training of large language and multimodal models.
  • Build scalable systems for rollout generation, inference, reward computation, and policy updates.
  • Improve distributed training efficiency, reliability, fault tolerance, and resource utilization.
  • Design interfaces that enable researchers to implement and evaluate new RL algorithms quickly.
  • Build tools for experiment configuration, checkpointing, logging, monitoring, and reproducibility.
  • Profile end-to-end training pipelines and resolve performance, memory, and communication bottlenecks.
  • Support on-policy and off-policy training workflows using verifiable, preference-based, or model-generated feedback.
  • Collaborate with researchers to translate experimental requirements into production-quality infrastructure.
  • Document system designs and contribute to technical reports and publications.

Skills

Python
PyTorch
Reinforcement Learning
Distributed Training
Multi-GPU Workloads
Experiment Tracking
Containerization

Education

PhD in Computer Science / ML / AI / Computer Engineering

Tools

CUDA
Docker
Kubernetes

Job description

ADVANCE YOUR CAREER. ADVANCE THE WORLD.

At AMD, we believetechnology has the power to solve the world's most important challenges. From advancing healthcare and scientific discovery to powering AI and the technologies people rely on every day, innovation at AMDis shapingthefuture.


Whetheryou'redesigning next-gen processors, enabling AI breakthroughs, orbringing leading edge products to market, every role at AMD contributes to something bigger- technologythat moves the world forward.Join us and, together, we'll advance your career.


JOB DETAILS:


  • Location: Santa Clara, CA, USA

  • Onsite/Hybrid: This role requires the student to work full time (40 hours a week), in either a hybrid or onsite work structure throughout the duration of the co-op/intern term

  • Duration: Summer 2027 Internship

    • Semester Schools: May 24, 2027 - August 13, 2027

    • Quarter Schools: June 21, 2027 - September 10, 2027




WHAT YOU WILL BE DOING:

We are seeking highly motivated AI Research intern to join our team. In this role -



  • Develop and optimize infrastructure for RL-based post-training of large language and multimodal models.

  • Build scalable systems for rollout generation, inference, reward computation, and policy updates.

  • Improve distributed training efficiency, reliability, fault tolerance, and resource utilization.

  • Design interfaces that enable researchers to implement and evaluate new RL algorithms quickly.

  • Build tools for experiment configuration, checkpointing, logging, monitoring, and reproducibility.

  • Profile end-to-end training pipelines and resolve performance, memory, and communication bottlenecks.

  • Support on-policy and off-policy training workflows using verifiable, preference-based, or model-generated feedback.

  • Collaborate with researchers to translate experimental requirements into production-quality infrastructure.

  • Document system designs and contribute to technical reports and publications.


WHO WE ARE LOOKING FOR:


  • Must be currently pursuing a PhD in Computer Science, Machine Learning, Artificial Intelligence, Computer Engineering, or a related field.

  • Strong programming skills in Python and experience with PyTorch.

  • Knowledge of reinforcement learning, LLM post-training, RLHF/RLAIF, or preference optimization.

  • Experience with distributed training, multi-GPU workloads, or large-scale inference.

  • Familiarity with parallelism strategies such as data, tensor, pipeline, or sequence parallelism.

  • Experience with training and inference frameworks, orchestration systems, or cluster environments.

  • Understanding of GPU performance, memory management, networking, and distributed communication.

  • Experience building reliable research infrastructure, profiling systems, or debugging distributed workloads.

  • Familiarity with containerization, experiment tracking, and cloud or cluster computing is beneficial.

  • Publications at leading venues such as ICML, NeurIPS, ICLR, MLSys, CVPR, ICCV, or ECCV are preferred.


Benefits offered are described: AMD benefits at a glance.


AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants' needs under the respective laws throughout all stages of the recruitment and selection process.


AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD's "Responsible AI Policy" is available here.


This posting is for an existing vacancy.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Summer 2027 PhD AI Research Infrastructure, RL Post-Training Intern
Summer 2027 PhD AI Research Infrastructure, RL Post-Training Intern

AMD • Santa Clara (CA)

Hybrid
USD 55,000 - 90,000
Summer 2027 PhD AI Research Infrastructure, RL Post-Training Intern
Summer 2027 PhD AI Research Infrastructure, RL Post-Training Intern

Advanced Micro Devices • Santa Clara (CA)

Hybrid
USD 50,000 - 67,000
Summer 2027 Master's AI Research, Reinforcement Learning and LLM Post-Training Intern
Summer 2027 Master's AI Research, Reinforcement Learning and LLM Post-Training Intern

AMD • Santa Clara (CA)

Hybrid
USD 36,000 - 47,000
Summer 2027 PhD Gen AI and Reinforcement Learning Research Intern
Summer 2027 PhD Gen AI and Reinforcement Learning Research Intern

AMD • Santa Clara (CA)

Hybrid
USD 34,000 - 55,000
Summer 2027 PhD Gen AI and Reinforcement Learning Research Intern
Summer 2027 PhD Gen AI and Reinforcement Learning Research Intern

Advanced Micro Devices • Santa Clara (CA)

Hybrid
USD 39,000 - 58,000
AMD benefits at a glance
Summer 2027 Master's AI Research, Reinforcement Learning and LLM Post-Training Intern
Summer 2027 Master's AI Research, Reinforcement Learning and LLM Post-Training Intern

Advanced Micro Devices • Santa Clara (CA)

On-site
USD 55,000 - 83,000
Summer 2027 PhD Gen AI and Reinforcement Learning Research Intern
Summer 2027 PhD Gen AI and Reinforcement Learning Research Intern

Advanced Micro Devices, Inc. • Santa Clara (CA)

Hybrid
USD 34,000 - 55,000
Summer 2027 PhD Technical Program Manager, AI Research Intern
Summer 2027 PhD Technical Program Manager, AI Research Intern

Advanced Micro Devices, Inc. • Santa Clara (CA)

Hybrid
USD 30,000 - 39,000
Summer 2027 PhD Technical Program Manager, AI Research Intern
Summer 2027 PhD Technical Program Manager, AI Research Intern

Advanced Micro Devices • Santa Clara (CA)

Hybrid
USD 34,000 - 48,000
Summer 2027 PhD Technical Program Manager, AI Research Intern
Summer 2027 PhD Technical Program Manager, AI Research Intern

AMD • Santa Clara (CA)

Hybrid
USD 20,000 - 33,000