ML Systems Research Engineer, RL / Inference / Agent Systems

AMD

Santa Clara (CA)

On-site

USD 180,000 - 250,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

AMD benefits

Job summary

AMD is hiring ML Systems Research Engineers to build reinforcement learning, inference, and evaluation infrastructure behind AI-for-engineering systems. This role focuses on scalable ML systems that support agents and models improving real engineering workflows, with tasks like running many attempts and measuring performance.

You will collaborate with scientists and engineers across compute optimization, verification, and tooling to make experiments reproducible and useful for production teams.

Qualifications

  • Bachelor’s degree in Computer Science, Computer Engineering, Electrical Engineering, Machine Learning, or related field, or equivalent practical experience. Master’s preferred; PhD is a plus, especially with work in ML systems, reinforcement learning, distributed systems, GPU computing, or AI infrastructure.
  • Experience building ML systems, RL infrastructure, inference services, agent frameworks, evaluation platforms, or distributed experimentation systems.
  • Strong understanding of model inference, batching, sampling, latency, throughput, observability, and reliability tradeoffs.
  • Ability to design experiments and evaluation pipelines with clear metrics, logs, reproducibility, and statistical discipline.
  • 1Strong collaboration skills with AI researchers, applied engineers, infrastructure engineers, and hardware domain experts.

Responsibilities

  • Build RL and inference systems for agentic engineering workflows, including job orchestration, sampling, scoring, caching, experiment tracking, and reproducible evaluation.
  • Develop infrastructure for long-horizon and high-latency reward tasks where validation can take minutes to hours.
  • Design staged rewards, proxy graders, sliced evaluation paths, retry strategies, and uncertainty-aware evaluation methods.
  • Support optimization workflows with systems for candidate generation, benchmark execution, correctness checking, profiler feedback, reward modeling, and model-level improvement.
  • Partner with AI research scientists on reward hacking research, reward shaping, metareasoning, and post-training methods for engineering tasks.
  • Build scalable inference and tool-use pipelines for LLM agents that interact with compilers, profilers, simulators, formal tools, benchmark harnesses, and internal knowledge sources.
  • Standardize datasets, eval definitions, run logs, leaderboards, failure taxonomies, and data collection for future training.
  • Analyze experimental results and turn system behavior into actionable guidance for model, agent, tool, and reward improvements.

Skills

Python programming
ML frameworks (PyTorch/JAX/TensorFlow)
RL infrastructure
Distributed experimentation
Observability and reliability
Experiment design and statistics
Collaboration with researchers/enginee

Education

Bachelor’s degree in CS/CE/EE/ML, or related; Master’s preferred

Tools

Kubernetes
Ray
Slurm
Workflow engines
Data pipelines
GPU profiling tools

Job description

ADVANCE YOUR CAREER. ADVANCE THE WORLD.

At AMD, we believe technology can change lives for the better. It can heal us, entertain us, and make us more connected, productive, and understanding of the world around us. And we’re looking for talent who feel the same: people who want to leave the planet better than they found it, those who don’t shy away from humanity’s challenges but are determined to help solve them.

ADVANCE YOUR CAREER. ADVANCE THE WORLD.

At AMD, we believe technology can change lives for the better. It can heal us, entertain us, and make us more connected, productive, and understanding of the world around us. And we’re looking for talent who feel the same: people who want to leave the planet better than they found it, those who don’t shy away from humanity’s challenges but are determined to help solve them.

AMD is powering the next generation of supercomputing, high-performance computing, cloud, and AI. Whether you’re designing next-gen processors, enabling AI breakthroughs, or creating go-to-market plans, every role at AMD contributes to something bigger — technology that moves the world forward.

THE ROLE

We are hiring ML Systems Research Engineers to build the reinforcement learning, inference, and evaluation infrastructure behind AI-for-engineering systems. This role focuses on the systems that let agents and models improve real engineering workflows: running many attempts, evaluating correctness, measuring performance, managing long-latency rewards, and feeding results back into model and agent improvement.

You will work across compute optimization, hardware engineering automation, verification, simulation, debugging. The emphasis is on scalable ML systems that make research practical, repeatable, and useful for production engineering teams.

THE PERSON

You are a systems-minded ML engineer or researcher who understands that model quality depends on the surrounding loop: data, tools, inference, graders, reward design, logging, and iteration speed. You can build reliable infrastructure, reason about RL and inference tradeoffs, and collaborate with scientists and applied engineers to make experiments reproducible and useful.

Key Responsibilities
  • Build RL and inference systems for agentic engineering workflows, including job orchestration, sampling, scoring, caching, experiment tracking, and reproducible evaluation.
  • Develop infrastructure for long-horizon and high-latency reward tasks where validation can take minutes to hours.
  • Design staged rewards, proxy graders, sliced evaluation paths, retry strategies, and uncertainty-aware evaluation methods.
  • Support optimization workflows with systems for candidate generation, benchmark execution, correctness checking, profiler feedback, reward modeling, and model-level improvement.
  • Partner with AI research scientists on reward hacking research, reward shaping, metareasoning, and post-training methods for engineering tasks.
  • Build scalable inference and tool-use pipelines for LLM agents that interact with compilers, profilers, simulators, formal tools, benchmark harnesses, and internal knowledge sources.
  • Standardize datasets, eval definitions, run logs, leaderboards, failure taxonomies, and data collection for future training.
  • Analyze experimental results and turn system behavior into actionable guidance for model, agent, tool, and reward improvements.
TECHNICAL FOCUS AREAS
  • Reinforcement learning and post-training infrastructure for tool-using agents.
  • Inference systems for LLMs and agents, including latency, throughput, batching, sampling, reliability, and observability.
  • Evaluation systems for tasks with expensive, delayed, mixed, or sparse rewards.
  • Reward design for engineering domains where correctness, performance, quality, and resource usage must be balanced.
  • Distributed experimentation, job orchestration, caching, data pipelines, dashboards, and reproducible run management.
  • Integration with external tools such as compilers, profilers, simulators, validation systems, benchmark harnesses, and ticketing or knowledge systems.
Preferred Qualifications
  • Strong programming skills in Python and experience with ML frameworks such as PyTorch, JAX, TensorFlow, or similar.
  • Experience building ML systems, RL infrastructure, inference services, agent frameworks, evaluation platforms, or distributed experimentation systems.
  • Strong understanding of model inference, batching, sampling, latency, throughput, observability, and reliability tradeoffs.
  • Ability to design experiments and evaluation pipelines with clear metrics, logs, reproducibility, and statistical discipline.
  • 1Strong collaboration skills with AI researchers, applied engineers, infrastructure engineers, and hardware domain experts.
Preferred Experience
  • Experience with reinforcement learning, RLHF, GRPO, preference optimization, reward modeling, reward shaping, or post-training systems.
  • Experience with LLM agents, tool-use systems, code generation, automated program repair, compiler optimization, or benchmark-driven development.
  • Experience with distributed systems, job orchestration, Kubernetes, Ray, Slurm, workflow engines, data pipelines, or large-scale experiment management.
  • Familiarity with GPU systems, ROCm/HIP, CUDA, profiling, kernel benchmarking, model serving, or distributed training/inference.
  • Exposure to hardware engineering workflows such as design, verification, firmware, simulation, or performance analysis is a strong plus.
  • Publications or shipped systems in ML systems, RL, inference optimization, AI infrastructure, or hardware/software co-design are valued.
EDUCATION

Bachelor’s degree in Computer Science, Computer Engineering, Electrical Engineering, Machine Learning, or related field, or equivalent practical experience. Master’s preferred; PhD is a plus, especially with work in ML systems, reinforcement learning, distributed systems, GPU computing, or AI infrastructure.

LOCATION

Santa Clara, CA

Benefits offered are described: AMD benefits at a glance.

AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee‑based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third‑party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process.

AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here.

This posting is for an existing vacancy.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Systems Research Engineer, RL / Inference / Agent Systems
ML Systems Research Engineer, RL / Inference / Agent Systems

Advanced Micro Devices • Santa Clara (CA)

Hybrid
USD 160,000 - 210,000
AMD benefits
AI Research Scientist, Reinforcement Learning (LLM) and Post-Training
AI Research Scientist, Reinforcement Learning (LLM) and Post-Training

AMD • Santa Clara (CA)

On-site
USD 180,000 - 260,000
AMD benefits
Gen AI Software Development Engineer
Gen AI Software Development Engineer

Jobzhr • Santa Clara (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
AI Research Scientist, Hardware AI Systems
AI Research Scientist, Hardware AI Systems

AMD • Santa Clara (CA)

On-site
USD 180,000 - 260,000
AMD benefits at a glance
Lead AI Infrastructure Engineer, Reinforcement Learning
Lead AI Infrastructure Engineer, Reinforcement Learning

Advanced Micro Devices • Santa Clara (CA)

On-site
USD 120,000 - 170,000
Lead AI Research Scientist, Hardware AI Systems
Lead AI Research Scientist, Hardware AI Systems

Advanced Micro Devices • Santa Clara (CA)

On-site
USD 130,000 - 160,000
Comprehensive benefits package
Diversity and inclusion initiatives
Sr. AI/ML Platform Engineer
Sr. AI/ML Platform Engineer

Advanced Micro Devices • Santa Clara (CA)

On-site
USD 190,000 - 260,000
Post-Training Platform Infrastructure Engineer
Post-Training Platform Infrastructure Engineer

Advanced Micro Devices • San Jose (CA)

Hybrid
USD 100,000 - 150,000
Comprehensive benefits
Collaborative work environment
Opportunities for career advancement
Lead AI Research Scientist - Infrastructure Engineer, Reinforcement Learning
Lead AI Research Scientist - Infrastructure Engineer, Reinforcement Learning

AMD • Santa Clara (CA)

On-site
USD 130,000 - 180,000
Competitive benefits package
ML and AI Knowledge Systems Engineer
ML and AI Knowledge Systems Engineer

Socket.dev • San Jose (CA)

Hybrid
USD 150,000 - 190,000