Applied AI/ML Engineer

Boundless

Northern (KY)

Hybrid

USD 175,000 - 250,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Competitive salary + equity
Health/dental/vision
Flexible PTO
Dev & conference budget
Remote-first with off-sites

Job summary

Boundless is building toward AI leadership by delivering end-to-end AI features on a growing GPU inference fleet. You will own AI products from prototype to production, optimizing LLM inference, and implementing RL/post-training harnesses to ensure stability and efficiency.

You’ll work with vLLM, SGLang, and Megatron-based tools, focusing on throughput, latency, and cost per token while operating with autonomy in a fast-moving environment.

Qualifications

  • 3+ years shipping ML/AI systems to production.
  • Experience with LLM inference using vLLM, SGLang, or TensorRT-LLM.
  • Experience with RL/post-training methods (GRPO, PPO, DPO, or SFT).
  • Strong Python and PyTorch development skills.
  • Understanding of GPU execution: batching, memory, CUDA basics.
  • Ability to operate with autonomy and bias for action.

Responsibilities

  • Own AI features from prototype to production — model selection, serving, evaluation, and iteration.
  • Deploy and optimize LLM inference across GPU fleet.
  • Build RL and post-training harnesses using slime (Megatron-LM + SGLang) and Prime Intellect (prime-rl + the Environments Hub / verifiers).
  • Build eval harnesses and benchmarks measuring quality, throughput, and cost.
  • Collaborate with Infrastructure on GPU scheduling and fleet utilization, and with Product on what to build next.

Skills

ML systems production
LLM inference experience
Python
PyTorch
CUDA basics
Ambiguity tolerance
LLM inference stacks
RL / post-training methods
Distributed training

Tools

Kubernetes
Container deployment
Ray
Slurm

Job description

Boundless is coordinating GPU compute at scale and building toward becoming a leader in AI. As an Applied AI/ML Engineer, you'll ship AI-powered products end-to-end on top of our growing GPU inference fleet — owning everything from serving low-latency inference to standing up reinforcement-learning post-training pipelines. This is a builder's role: you take an idea from prototype to production, tune it for throughput and cost on real GPUs, and iterate fast on customer and internal feedback.

You should be comfortable operating with a high degree of autonomy, navigating ambiguity, and defaulting to a strong bias for action.

What You'll Do
End-to-End AI Product Delivery

Own AI features and products from prototype through production — model selection, serving, evaluation, and iteration — shipping working software rather than research artifacts.

Inference Serving

Deploy and optimize LLM inference across the fleet using vLLM and SGLang. Tune continuous batching, KV-cache management, quantization, speculative decoding, and multi-model routing to maximize throughput and minimize latency and cost per token.

RL & Post-Training Harnesses

Build and operate reinforcement-learning and post-training pipelines using slime (Megatron-LM + SGLang) and Prime Intellect (prime-rl + the Environments Hub / verifiers). This includes reward and verifier design, rollout orchestration, weight synchronization, and keeping long-running training stable.

Evaluation & Iteration

Build eval harnesses and benchmarks that measure quality, throughput, and cost together, and use them to drive fast, data-informed iteration.

Work Across the Stack

Partner with Infrastructure on GPU scheduling and fleet utilization, and with Product on what to build next and why.

  • 3+ years shipping ML/AI systems to production
  • Hands-on experience serving LLM inference with vLLM, SGLang, or TensorRT-LLM
  • Experience with RL / post-training methods (GRPO, PPO, DPO, or SFT), or strong adjacent experience and a clear desire to go deep here
  • Strong Python and PyTorch
  • Working understanding of GPU execution: batching, memory, and basic CUDA concepts
  • Comfort operating in ambiguity with a strong bias for action
Nice to Have
  • Direct experience with slime, prime-rl, the verifiers library, or Megatron-LM
  • Distributed training experience (FSDP, TP/PP/DP parallelism)
  • Quantization (FP8/INT8), P/D disaggregation, or speculative decoding
  • Experience with verifiable inference or large-scale distributed systems
  • Kubernetes and container-based deployment
  • Familiarity with GPU fleet orchestration (Ray, SkyPilot, Slurm)
Additional Requirements
  • Candidates must include a public GitHub profile in their application.
  • The GitHub profile should demonstrate a minimum of 1 year of activity/history.
  • Applications that do not include a GitHub profile, or show insufficient activity, will not be considered.

At Boundless, we take care of our people, because building the future of AI compute starts with an empowered team. Here's what you can expect when you join us:

  • Competitive salary (proposed band b/t US$175k and $250k annually) + equity allocation
  • Health, dental, vision (for U.S. employees; region-adjusted globally)
  • Flexible PTO
  • Professional development and conference travel budget
  • Remote-first with regular off-sites and a high-trust, high-velocity team environment

We are a global team, and applicants from around the world are welcome to apply.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Infrastructure Engineer - GPU Compute
Senior Infrastructure Engineer - GPU Compute

Boundless • Northern (KY)

Hybrid
USD 175,000 - 250,000
Equity allocation
Remote-first with off-sites
Health, dental, vision
Founding Product Manager
Founding Product Manager

Boundless • Northern (KY)

Hybrid
USD 100,000 - 150,000
Equity
Health, dental, vision
Flexible PTO
+2
Technical Business Developer
Technical Business Developer

Boundless • San Francisco (CA), Northern (KY)

Hybrid
USD 100,000 - 150,000
Equity
Health insurance
Remote-friendly
+1
Applied AI/ML Engineer — Build AI at Scale (Remote)
Applied AI/ML Engineer — Build AI at Scale (Remote)

Boundless Networks • United States

Remote
USD 175,000 - 250,000
Equity allocation
Health, dental, vision
Flexible PTO
+3
Remote AI/ML Engineer: Production AI & RL Pipelines
Remote AI/ML Engineer: Production AI & RL Pipelines

Boundless Networks, Inc. • United States

On-site
USD 175,000 - 250,000
Equity allocation
Remote-first
Off-sites and team events
+1
Member of Technical Staff - Inference
Member of Technical Staff - Inference

Prime Intellect • United States

Hybrid
USD 120,000 - 150,000
Competitive compensation
Flexible work arrangement
Full visa sponsorship
+2
Member of Technical Staff, Post-Training & Applied Research
Member of Technical Staff, Post-Training & Applied Research

San Francisco Tensor Company • San Francisco (CA)

On-site
USD 275,000 - 315,000
Relocation assistance
Senior AI Infra & GPU Compute Product Manager
Senior AI Infra & GPU Compute Product Manager

Boundless • United States

On-site
USD 100,000 - 150,000
Equity
Health insurance
Flexible PTO
+2
Remote AI/ML Engineer — Production Inference & RL
Remote AI/ML Engineer — Production Inference & RL

Boundless • Northern (KY)

Hybrid
USD 175,000 - 250,000
Competitive salary + equity
Health/dental/vision
Flexible PTO
+2
Inference Performance Engineer
Inference Performance Engineer

Adaption Labs • San Francisco (CA)

On-site
USD 180,000 - 260,000
Flexible work
Adaption Passport
Lunch stipend
+1