Remote AI/ML Engineer — Production Inference & RL

Boundless

Northern (KY)

Hybrid

USD 175,000 - 250,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Competitive salary + equity
Health/dental/vision
Flexible PTO
Dev & conference budget
Remote-first with off-sites

Job summary

Boundless is building toward AI leadership by delivering end-to-end AI features on a growing GPU inference fleet. You will own AI products from prototype to production, optimizing LLM inference, and implementing RL/post-training harnesses to ensure stability and efficiency.

You’ll work with vLLM, SGLang, and Megatron-based tools, focusing on throughput, latency, and cost per token while operating with autonomy in a fast-moving environment.

Qualifications

  • 3+ years shipping ML/AI systems to production.
  • Experience with LLM inference using vLLM, SGLang, or TensorRT-LLM.
  • Experience with RL/post-training methods (GRPO, PPO, DPO, or SFT).
  • Strong Python and PyTorch development skills.
  • Understanding of GPU execution: batching, memory, CUDA basics.
  • Ability to operate with autonomy and bias for action.

Responsibilities

  • Own AI features from prototype to production — model selection, serving, evaluation, and iteration.
  • Deploy and optimize LLM inference across GPU fleet.
  • Build RL and post-training harnesses using slime (Megatron-LM + SGLang) and Prime Intellect (prime-rl + the Environments Hub / verifiers).
  • Build eval harnesses and benchmarks measuring quality, throughput, and cost.
  • Collaborate with Infrastructure on GPU scheduling and fleet utilization, and with Product on what to build next.

Skills

ML systems production
LLM inference experience
Python
PyTorch
CUDA basics
Ambiguity tolerance
LLM inference stacks
RL / post-training methods
Distributed training

Tools

Kubernetes
Container deployment
Ray
Slurm

Job description

Boundless is building toward AI leadership by delivering end-to-end AI features on a growing GPU inference fleet. You will own AI products from prototype to production, optimizing LLM inference, and implementing RL/post-training harnesses to ensure stability and efficiency.

You’ll work with vLLM, SGLang, and Megatron-based tools, focusing on throughput, latency, and cost per token while operating with autonomy in a fast-moving environment.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Remote AI/ML Engineer: Production AI & RL Pipelines
Remote AI/ML Engineer: Production AI & RL Pipelines

Boundless Networks, Inc. • United States

On-site
USD 175,000 - 250,000
Equity allocation
Remote-first
Off-sites and team events
+1
Applied AI/ML Engineer — Build AI at Scale (Remote)
Applied AI/ML Engineer — Build AI at Scale (Remote)

Boundless Networks • United States

Remote
USD 175,000 - 250,000
Equity allocation
Health, dental, vision
Flexible PTO
+3
Applied AI/ML Engineer
Applied AI/ML Engineer

Boundless • Northern (KY)

Hybrid
USD 175,000 - 250,000
Competitive salary + equity
Health/dental/vision
Flexible PTO
+2
Senior AI Infra & GPU Compute Product Manager
Senior AI Infra & GPU Compute Product Manager

Boundless • United States

On-site
USD 100,000 - 150,000
Equity
Health insurance
Flexible PTO
+2
Staff Engineer, RL Inference & Distributed Systems
Staff Engineer, RL Inference & Distributed Systems

Pantera Capital • Palo Alto (CA)

On-site
USD 150,000 - 230,000
Senior AI Engineer: Production LLM Agents & Inference
Senior AI Engineer: Production LLM Agents & Inference

Bitus Labs • Irvine (CA)

Hybrid
USD 140,000 - 190,000
Senior Inference & RL Systems Engineer (Scalable ML Infra)
Senior Inference & RL Systems Engineer (Scalable ML Infra)

Magic AI, Inc • San Francisco (CA)

On-site
USD 300,000 - 550,000
Equity compensation
401(k) matching
Health, dental and vision insurance
+4
Member of Technical Staff - RL Inference
Member of Technical Staff - RL Inference

Pantera Capital • Palo Alto (CA)

On-site
USD 150,000 - 230,000
Senior AI Engineer – LLM Agents & Inference (Mandarin Required)
Senior AI Engineer – LLM Agents & Inference (Mandarin Required)

Bitus Labs • Irvine (CA)

Hybrid
USD 140,000 - 190,000
Staff Software Engineer, AI Inference
Staff Software Engineer, AI Inference

ChatGPT Jobs • New York (NY)

On-site
USD 180,000 - 240,000
Health Insurance
Equity