Principal LLM Inference Engineer — AI Platform

JPMorganChase

Jersey City (NJ)

On-site

USD 210,000 - 320,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

JPMorganChase seeks a Principal Software Engineer to own LLM inference performance, benchmarking, and efficiency at scale within the AI/ML Data Platform. You will shape optimization strategy, run rigorous experiments, and guide production readiness for fast, cost-efficient model serving.

You will collaborate with EKS and disaggregated serving teams, drive progressive decoding strategies, and influence leadership on safe scaling practices.

Qualifications

  • 7+ years of software engineering experience.
  • Hands-on experience with LLM inference systems such as vLLM, TensorRT-LLM, SGLang, LLM-D.
  • Strong understanding of GPU memory architecture and KV cache dynamics.
  • Experience with quantization techniques at scale.
  • Familiarity with speculative decoding and drivers of acceptance rates.
  • Rigorous benchmarking with numbers behind every claim.
  • Comfort operating in cloud GPU infrastructure (AWS; EKS).
  • Ability to communicate trade-offs to senior stakeholders.

Responsibilities

  • Own systematic benchmarking and performance characterization across all production LLM workloads.
  • Design and execute quantization experiments — FP8, INT8/INT4, measure accuracy delta and throughput.
  • Drive speculative decoding strategy across the model portfolio and optimize acceptance rate.
  • Build and maintain a GPU efficiency scorecard for leadership visibility.
  • Benchmark our platform vs external providers and industry numbers.
  • Lead inference engine upgrade evaluations with new schedulers and async tensor parallelism.
  • Collaborate with KV-cache optimization and multi-node serving architecture teams.
  • Design and run GPU chaos engineering and failure scenario testing.
  • Architect agentic AI-enabled development workflows with guardrails for validation and security.

Skills

LLM Inference Systems
GPU Optimization
Benchmarking & Performance
Quantization techniques
Speculative decoding
Cloud GPUs (AWS/EKS)

Job description

JPMorganChase seeks a Principal Software Engineer to own LLM inference performance, benchmarking, and efficiency at scale within the AI/ML Data Platform. You will shape optimization strategy, run rigorous experiments, and guide production readiness for fast, cost-efficient model serving.

You will collaborate with EKS and disaggregated serving teams, drive progressive decoding strategies, and influence leadership on safe scaling practices.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior LLM Inference Engineer – AI Platform Lead
Senior LLM Inference Engineer – AI Platform Lead

Fairygodboss • Jersey City (NJ)

On-site
USD 190,000 - 240,000
Principal AI Engineer - Lead LLM Platform & Innovation
Principal AI Engineer - Lead LLM Platform & Innovation

Next Frontier Capital • Palo Alto (CA)

On-site
USD 210,000 - 290,000
Senior Principal LLM Engineer — AI Platform & Scale
Senior Principal LLM Engineer — AI Platform & Scale

JPMorgan Chase & Co. • Palo Alto (CA)

On-site
USD 233,000 - 325,000
GenAI Platform Architect for High-Scale LLM Inference
GenAI Platform Architect for High-Scale LLM Inference

Socket.dev • New Jersey

On-site
USD 180,000 - 240,000
Principal Software Engineer - LLM Optimization
Principal Software Engineer - LLM Optimization

JPMorgan Chase & Co. • Jersey City (NJ)

On-site
USD 180,000 - 240,000
Senior Principal AI/ML Engineer — LLM & GNN Architect
Senior Principal AI/ML Engineer — LLM & GNN Architect

Fairygodboss • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Principal Software Engineer - LLM Optimization
Principal Software Engineer - LLM Optimization

JPMorganChase • Jersey City (NJ)

On-site
USD 210,000 - 320,000
Principal Software Engineer - LLM Optimization
Principal Software Engineer - LLM Optimization

Fairygodboss • Jersey City (NJ)

On-site
USD 190,000 - 240,000
Lead AI Engineer for Enterprise LLM Platform
Lead AI Engineer for Enterprise LLM Platform

Next Frontier Capital • Palo Alto (CA)

On-site
USD 180,000 - 260,000
Senior AI Platform Engineer: ML/LLM Ops
Senior AI Platform Engineer: ML/LLM Ops

JPMorgan Chase & Co. • Jersey City (NJ)

On-site
USD 180,000 - 240,000