AI Inference Platform Engineer

DRW

Chicago (IL)

On-site

USD 200,000 - 250,000

Full time

6 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Group medical
Dental and vision insurance
401k
Disability insurance
Life/AD&D insurance
Health Savings Account
Flexible Spending Accounts

Job summary

DRW is seeking an AI Inference Platform Engineer to build, operate, and optimize the systems that serve large language, vision, multimodal, and embedding models across DRW. You’ll own the end-to-end serving platform from onboarding models to performance, reliability, and cost optimization in a production environment.

You will optimize GPU-backed inference, design distributed architectures, and collaborate with SRE to automate deployment and observability.

Qualifications

  • Hands-on experience serving LLMs on NVIDIA GPUs.
  • Familiarity with modern inference runtimes and architectures.
  • Experience with model serving, CI/CD, and observability.

Responsibilities

  • Optimize LLM inference performance across NVIDIA GPUs and runtimes.
  • Build end-to-end profiling and observability to locate bottlenecks.
  • Design KV cache and distributed inference architectures.
  • Onboard new models and configure runtimes, precision, sharding, and batching.
  • Maintain performance profiles and run regression tests.
  • Measure quality equivalence across serving configurations.
  • Manage production serving lifecycle: versioning, staging, canarying, rollback.
  • Collaborate with SRE to automate deployment and reliability.
  • Optimize placement, scaling, and resource allocation for cost efficiency.
  • Design multi-tenant scheduling and isolation across shared GPU capacity.

Skills

GPU inference
Performance profiling
Multi-tenant scheduling
Systems troubleshooting
Linux/Unix

Tools

TensorRT-LLM
vLLM
SGLang

Job description

DRW is a diversified trading firm with over 3 decades of experience bringing sophisticated technology and exceptional people together to operate in markets around the world. We value autonomy and the ability to quickly pivot to capture opportunities, so we operate using our own capital and trading at our own risk. Headquartered in Chicago with offices throughout the U.S., Canada, Europe, and Asia, we trade a variety of asset classes including Fixed Income, ETFs, Equities, FX, Commodities and Energy across all major global markets. We have also leveraged our expertise and technology to expand into three non-traditional strategies: real estate, venture capital and cryptoassets. We operate with respect, curiosity and open minds. The people who thrive here share our belief that it’s not just what we do that matters–it's how we do it. DRW is a place of high expectations, integrity, innovation and a willingness to challenge consensus.

About the Role

We're looking for an AI Inference Platform Engineer to build, operate, and optimize the systems that serve large language, vision, multimodal, and embedding models across DRW. This role provides DRW's firmwide interface to modern AI models, from early evaluation through reliable production use.

You’ll work across inference runtimes, distributed systems, and production platform engineering, with deep GPU literacy. You’ll own the serving platform end-to-end: onboarding newly released models, measuring quality and performance equivalence across serving configurations, scheduling workloads across tenants, and continuously improving latency, throughput, utilization, reliability, and cost across the inference fleet.

What You’ll Do
  • Optimize LLM inference performance across modern NVIDIA GPU architectures and inference runtimes.
  • Build end-to-end performance profiling and observability to identify bottlenecks from individual GPU kernels through multi-node inference systems.
  • Design and optimize KV cache and distributed inference architectures, including caching, routing, memory tiering, and prefill/decode strategies.
  • Own day-0 model onboarding, determining the appropriate runtime, precision, sharding, memory, batching, cache policy, and serving configuration for new models.
  • Maintain validated performance profiles for important model and hardware combinations, including performance and quality regression testing.
  • Measure and monitor quality equivalence across serving configurations, including KV cache quantization, speculative decoding acceptance thresholds, precision choices, and model routing, so in-house serving can be trusted to match reference-model quality on production workloads.
  • Manage the production serving lifecycle of models, including versioning, compatibility, staging, canarying, promotion, rollback, and retirement.
  • Partner with SRE and platform teams to automate model deployment, distribution, production readiness, observability, and reliable operation across environments.
  • Optimize model placement, scaling, and resource allocation across the inference fleet to improve utilization and cost efficiency while meeting performance and reliability requirements.
  • Design and operate multi-tenant scheduling and isolation across shared GPU capacity, balancing latency SLOs, throughput, and priority across concurrent workloads.
What We’re Looking For
The Tech
  • Hands‑on experience serving LLMs on NVIDIA GPUs, with familiarity across current and emerging architectures (Hopper, Blackwell, and successors), HBM, Tensor Cores, NVLink/NVSwitch, and the compute and memory bottlenecks that shape serving decisions.
  • Deep expertise in at least one modern inference runtime such as TensorRT-LLM, vLLM, or SGLang.
  • Practical knowledge of inference optimization techniques including continuous batching, scheduling, chunked prefill, speculative decoding, quantization, CUDA Graphs, and paged attention.
  • Understanding of KV cache architecture, including prefix caching, block management, sizing, eviction, quantization, cache‑aware routing, and multi‑tier caching.
  • Experience measuring model quality equivalence across serving configurations, including evaluation harnesses, task‑specific benchmarks, and regression detection for quantization, KV cache, and speculative decoding changes.
  • Experience designing and tuning distributed inference systems, including tensor parallelism, multi‑node deployments, and disaggregated prefill and decode.
  • Experience with multi‑tenant GPU scheduling, workload isolation, and QoS across concurrent inference workloads.
  • Proficiency with GPU performance and observability tooling such as Nsight, DCGM, OpenTelemetry, Prometheus, and Grafana.
  • Strong Linux and systems performance fundamentals, with the ability to diagnose bottlenecks across hardware, drivers, runtimes, networking, and application layers.
  • Production experience with model serving infrastructure, including CI/CD, automated testing, observability, and production readiness.
The Intangibles
  • You take a measurement‑driven approach to performance optimization.
  • You take ownership of performance problems across hardware, runtime, model, and infrastructure boundaries.
  • You can move quickly and reprioritize as trading needs change, while maintaining a high bar for production systems.
  • You understand the importance of reliability, predictability, and performance when AI systems are integrated into trading workflows and decision‑making processes.
  • You can evaluate unfamiliar models, runtimes, and hardware quickly and make sound engineering decisions with limited prior guidance.
  • You communicate clearly and can explain complex performance tradeoffs across engineering teams.
Compensation and Benefits

The annual base salary range for this position is $200,000 to $250,000 depending on the candidate’s experience, qualifications, and relevant skill set. The position is also eligible for an annual discretionary bonus.

DRW offers a comprehensive suite of employee benefits including:

  • group medical
  • pharmacy
  • dental and vision insurance
  • 401k (with discretionary employer match)
  • short and long‑term disability
  • life and AD&D insurance
  • health savings accounts
  • flexible spending accounts

For more information about DRW's processing activities and our use of job applicants' data, please view our Privacy Notice at https://drw.com/privacy-notice. California residents, please review the California Privacy Notice for information about certain legal rights at https://drw.com/california-privacy-notice.

[#LI-VD1]

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Inference Platform Engineer
AI Inference Platform Engineer

DRW Holdings, LLC • Chicago (IL)

On-site
USD 200,000 - 250,000
Medical Insurance
Dental Insurance
Vision Insurance
+5
AI Inference Platform Engineer DRW · · Chicago, United States 18 hours ago
AI Inference Platform Engineer DRW · · Chicago, United States 18 hours ago

Tradermath • Chicago (IL), Northern (KY)

Hybrid
USD 200,000 - 250,000
Medical insurance
Dental insurance
Vision insurance
+1
Platform Engineer - AI Engineering
Platform Engineer - AI Engineering

DRW • Chicago (IL)

On-site
USD 130,000 - 190,000
Health, dental and vision insurance
401k with discretionary employer match
HSA/FSAs and other benefits
Platform Engineer - AI Engineering
Platform Engineer - AI Engineering

DRW Holdings, LLC. • Chicago (IL)

On-site
USD 130,000 - 190,000
Group medical insurance
401k with employer match
Disability insurance
+3
High-Performance AI Inference Platform Engineer
High-Performance AI Inference Platform Engineer

DRW Holdings, LLC • Chicago (IL)

On-site
USD 200,000 - 250,000
Medical Insurance
Dental Insurance
Vision Insurance
+5
Team Lead, AI Engineering
Team Lead, AI Engineering

DRW Holdings, LLC. • Chicago (IL)

On-site
USD 200,000 - 250,000
Employee benefits
GPU‑Focused AI Inference Platform Engineer
GPU‑Focused AI Inference Platform Engineer

DRW • Chicago (IL)

On-site
USD 200,000 - 250,000
Group medical
Dental and vision insurance
401k
+4
Senior AI Inference Platform Engineer
Senior AI Inference Platform Engineer

Tradermath • Chicago (IL), Northern (KY)

Hybrid
USD 200,000 - 250,000
Medical insurance
Dental insurance
Vision insurance
+1
Quantitative AI Strategist
Quantitative AI Strategist

Trading Interview • New York (NY)

On-site
USD 175,000 - 250,000
Group medical insurance
Pharmacy insurance
Dental and vision insurance
+5
Team Lead, AI Engineering
Team Lead, AI Engineering

DRW • Chicago (IL)

On-site
USD 200,000 - 250,000
Medical, dental, and vision insurance
401k with discretionary employer match
Short and long-term disability
+3