Lead AI/ML Platform Engineer — LLM Inference & Scale

JPMorgan Chase & Co.

Auchentibber

On-site

GBP 90,000 - 140,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

JPMorganChase is building AI infrastructure to power enterprise-scale LLM inference. We seek a Lead Software Engineer to drive performance, quantization and efficiency across production workloads in a scalable AI platform.

You will benchmark, optimize, and validate inference engines, contribute to scheduling and speculative decoding strategies, and collaborate with senior engineers to deliver fast, cost-efficient, production-ready models at scale.

Qualifications

  • Formal training or certification on software engineering concepts and advanced applied experience – preferably Go / Python.
  • Hands-on experience with LLM inference systems — vLLM, TensorRT-LLM, SGLang, LLM-D, or equivalent production serving engines.
  • Strong understanding of GPU memory architecture, including KV cache sizing and dynamics, memory-bandwidth versus compute bottlenecks, and the practical implications of quantization at inference time.
  • Experience with quantization techniques and their real-world tradeoffs at scale.
  • Familiarity with speculative decoding and the variables that drive acceptance rates in production workloads.
  • Rigorous benchmarking skills using GuideLLM, custom harnesses, or equivalent tooling, with the ability to support every performance claim with data.
  • Experience operating in cloud GPU infrastructure at scale (AWS, Kubernetes-based managed inference services).
  • Ability to communicate technical trade-offs clearly to engineering peers and senior stakeholders.
  • Hands-on experience using enterprise-authorized AI-assisted software development tools within the work environment (e.g., for coding, testing, troubleshooting, or documentation) with demonstrated ability to critically evaluate and validate AI-generated outputs.
  • Understanding of responsible AI use in engineering workflows, including data sensitivity considerations, secure handling of inputs/outputs, and adherence to resiliency and security expectations.

Responsibilities

  • Execute systematic benchmarking and performance characterization across production LLM workloads, establishing reproducible baselines, identifying regressions, and quantifying the impact of configuration changes before they reach production.
  • Design and run quantization experiments — FP8, INT8/INT4 (GPTQ/AWQ), and next-generation precision formats — measuring accuracy delta, throughput improvement, memory reduction, and cost-per-token impact.
  • Support speculative decoding strategy across the model portfolio, including draft model, n-gram, and multi-token prediction approaches, contributing to acceptance rate measurement and per-workload configuration recommendations.
  • Build and maintain GPU efficiency metrics covering utilization, memory headroom, cost per 1K tokens, and waste identification — providing engineering teams with a data-driven view of platform efficiency.
  • Benchmark the platform against external providers and published industry numbers, identifying gaps and contributing to improvement initiatives.
  • Participate in inference engine upgrade evaluations, including new scheduler architectures, async tensor parallelism, disaggregated prefill/decode, and advanced speculative decoding, supporting systematic validation before production promotion.
  • Contribute to GPU chaos engineering efforts, including induced failure scenarios, hardware diagnostic monitoring, and detection and recovery measurement.
  • Leverage enterprise-authorized AI coding assist tools within the work environment to improve code quality, delivery speed, and productivity, while validating outputs through peer review, automated testing, and secure coding standards.
  • Apply knowledge of tools within the SDLC toolchain, to improve automation value realized by automation.

Skills

Go/Python
LLM inference
GPU performance
Benchmarking
Quantization
Cloud GPUs
Communication

Tools

vLLM
TensorRT-LLM
SGLang
LLM-D

Job description

JPMorganChase is building AI infrastructure to power enterprise-scale LLM inference. We seek a Lead Software Engineer to drive performance, quantization and efficiency across production workloads in a scalable AI platform.

You will benchmark, optimize, and validate inference engines, contribute to scheduling and speculative decoding strategies, and collaborate with senior engineers to deliver fast, cost-efficient, production-ready models at scale.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Platform Engineer — LLM Inference (Go/Python)
Senior AI Platform Engineer — LLM Inference (Go/Python)

J.P. MORGAN • Scotland

On-site
GBP 110,000 - 170,000
Senior AI/ML Platform Engineer – Scale & Inference
Senior AI/ML Platform Engineer – Scale & Inference

JPMorganChase • Greater London

On-site
GBP 90,000 - 115,000
Senior AI Platform Engineer – LLM Inference Backend
Senior AI Platform Engineer – LLM Inference Backend

JPMorgan Chase & Co. • Greater London

On-site
GBP 110,000 - 150,000
Lead AI/LLM Inference Engineer (Go/Python) for Production
Lead AI/LLM Inference Engineer (Go/Python) for Production

JPMorgan Chase & Co. • Glasgow

On-site
GBP 90,000 - 130,000
Lead Software Engineer - Python / Go & AI/ML
Lead Software Engineer - Python / Go & AI/ML

JPMorgan Chase & Co. • Auchentibber

On-site
GBP 90,000 - 140,000
Lead Software Engineer - Python / Go & AI/ML
Lead Software Engineer - Python / Go & AI/ML

JPMorgan Chase & Co. • Glasgow

On-site
GBP 90,000 - 130,000
Senior AI Infra Engineer – LLM Ops & Reliability
Senior AI Infra Engineer – LLM Ops & Reliability

JPMorgan Chase & Co. • Auchentibber

On-site
GBP 90,000 - 150,000
Lead, GenAI & LLM Platform Engineer
Lead, GenAI & LLM Platform Engineer

JPMorgan Chase & Co. • Greater London

On-site
GBP 80,000 - 100,000
Senior Lead AI/ML Platform Engineer
Senior Lead AI/ML Platform Engineer

JPMorganChase • Glasgow

On-site
GBP 90,000 - 130,000
Senior Lead, AI Infra Reliability Engineer
Senior Lead, AI Infra Reliability Engineer

JPMorganChase • Glasgow

On-site
GBP 90,000 - 130,000