RL & Inference ML Systems Engineer for Engineering AI

AMD

Santa Clara (CA)

On-site

USD 180,000 - 250,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

AMD benefits

Job summary

AMD is hiring ML Systems Research Engineers to build reinforcement learning, inference, and evaluation infrastructure behind AI-for-engineering systems. This role focuses on scalable ML systems that support agents and models improving real engineering workflows, with tasks like running many attempts and measuring performance.

You will collaborate with scientists and engineers across compute optimization, verification, and tooling to make experiments reproducible and useful for production teams.

Qualifications

  • Bachelor’s degree in Computer Science, Computer Engineering, Electrical Engineering, Machine Learning, or related field, or equivalent practical experience. Master’s preferred; PhD is a plus, especially with work in ML systems, reinforcement learning, distributed systems, GPU computing, or AI infrastructure.
  • Experience building ML systems, RL infrastructure, inference services, agent frameworks, evaluation platforms, or distributed experimentation systems.
  • Strong understanding of model inference, batching, sampling, latency, throughput, observability, and reliability tradeoffs.
  • Ability to design experiments and evaluation pipelines with clear metrics, logs, reproducibility, and statistical discipline.
  • 1Strong collaboration skills with AI researchers, applied engineers, infrastructure engineers, and hardware domain experts.

Responsibilities

  • Build RL and inference systems for agentic engineering workflows, including job orchestration, sampling, scoring, caching, experiment tracking, and reproducible evaluation.
  • Develop infrastructure for long-horizon and high-latency reward tasks where validation can take minutes to hours.
  • Design staged rewards, proxy graders, sliced evaluation paths, retry strategies, and uncertainty-aware evaluation methods.
  • Support optimization workflows with systems for candidate generation, benchmark execution, correctness checking, profiler feedback, reward modeling, and model-level improvement.
  • Partner with AI research scientists on reward hacking research, reward shaping, metareasoning, and post-training methods for engineering tasks.
  • Build scalable inference and tool-use pipelines for LLM agents that interact with compilers, profilers, simulators, formal tools, benchmark harnesses, and internal knowledge sources.
  • Standardize datasets, eval definitions, run logs, leaderboards, failure taxonomies, and data collection for future training.
  • Analyze experimental results and turn system behavior into actionable guidance for model, agent, tool, and reward improvements.

Skills

Python programming
ML frameworks (PyTorch/JAX/TensorFlow)
RL infrastructure
Distributed experimentation
Observability and reliability
Experiment design and statistics
Collaboration with researchers/enginee

Education

Bachelor’s degree in CS/CE/EE/ML, or related; Master’s preferred

Tools

Kubernetes
Ray
Slurm
Workflow engines
Data pipelines
GPU profiling tools

Job description

AMD is hiring ML Systems Research Engineers to build reinforcement learning, inference, and evaluation infrastructure behind AI-for-engineering systems. This role focuses on scalable ML systems that support agents and models improving real engineering workflows, with tasks like running many attempts and measuring performance.

You will collaborate with scientists and engineers across compute optimization, verification, and tooling to make experiments reproducible and useful for production teams.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Systems Engineer for RL & Inference Infrastructure
ML Systems Engineer for RL & Inference Infrastructure

Advanced Micro Devices • Santa Clara (CA)

Hybrid
USD 160,000 - 210,000
AMD benefits
ML Systems Research Engineer, RL / Inference / Agent Systems
ML Systems Research Engineer, RL / Inference / Agent Systems

Advanced Micro Devices • Santa Clara (CA)

Hybrid
USD 160,000 - 210,000
AMD benefits
ML Systems Research Engineer, RL / Inference / Agent Systems
ML Systems Research Engineer, RL / Inference / Agent Systems

AMD • Santa Clara (CA)

On-site
USD 180,000 - 250,000
AMD benefits
Lead RL Infra Engineer - Scalable GPU Training Platforms
Lead RL Infra Engineer - Scalable GPU Training Platforms

Advanced Micro Devices • Santa Clara (CA)

On-site
USD 120,000 - 170,000
Staff AI Systems Engineer — Inference & RL
Staff AI Systems Engineer — Inference & RL

Together • San Francisco (CA)

On-site
USD 200,000 - 280,000
Health insurance
Startup equity
Competitive benefits
AI Systems Engineer: ML Kernels & HPC Acceleration
AI Systems Engineer: ML Kernels & HPC Acceleration

Socket.dev • San Jose (CA)

Hybrid
USD 150,000 - 190,000
Lead RL Scientist for LLM Post-Training & Code Models
Lead RL Scientist for LLM Post-Training & Code Models

AMD • Santa Clara (CA)

On-site
USD 130,000 - 160,000
Comprehensive benefits package
RL Systems Engineer: Inference & Training at Scale
RL Systems Engineer: Inference & Training at Scale

xAI • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Post-Training LLM Inference Platform Engineer
Post-Training LLM Inference Platform Engineer

Advanced Micro Devices • San Jose (CA)

Hybrid
USD 100,000 - 150,000
Comprehensive benefits
Collaborative work environment
Opportunities for career advancement
Staff Engineer, RL Inference & Distributed Systems
Staff Engineer, RL Inference & Distributed Systems

Pantera Capital • Palo Alto (CA)

On-site
USD 150,000 - 230,000