Senior LLM Inference Engineer – AI Platform Lead

Fairygodboss

Jersey City (NJ)

On-site

USD 190,000 - 240,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

JPMorganChase is seeking a Principal Software Engineer to own LLM inference performance and optimization for scalable AI platforms. You will benchmark, quantify improvements, and drive decisions with leaders to improve speed, cost efficiency, and production readiness across models.

You will collaborate with EKS and disaggregated serving teams, lead quantization experiments, and shape guardrails for secure and auditable AI workflows.

Qualifications

  • Formal training or certification on software engineering concepts and 7+ years applied experience.
  • Deep, hands-on experience with LLM inference systems - vLLM, TensorRT-LLM, SGLang, LLM-D or equivalent production serving engines.
  • Strong grasp of GPU memory architecture and quantization at inference time.
  • Experience with quantization techniques and real-world tradeoffs at scale.
  • Familiarity with speculative decoding and acceptance-rate variables in production workloads.
  • Rigorous benchmarking instincts with data-backed claims.

Responsibilities

  • Own systematic benchmarking and performance characterization across all production LLM workloads.
  • Design and execute quantization experiments - FP8, INT8/INT4, measuring accuracy delta, throughput, memory, cost per token.
  • Drive speculative decoding strategy across the model portfolio and provide configuration recommendations.
  • Build and maintain a GPU efficiency scorecard for leadership review.
  • Benchmark platform against external providers and published industry numbers.
  • Lead inference engine upgrade evaluations and validate before production promotion.
  • Collaborate on KV-cache optimization, prefix caching, and multi-node serving architecture.
  • Design and run GPU chaos engineering, monitoring, and recovery measurement.
  • Architect AI-enabled engineering workflows with guardrails for validation, security, resiliency.
  • Apply SDLC toolchain and enterprise AI capabilities to improve automation and quality.

Skills

LLM inference systems
GPU memory architecture
quantization techniques
speculative decoding
benchmarking
cloud GPU infrastructure
stakeholder communication
security/resiliency governance

Tools

vLLM
TensorRT-LLM
SGLang
LLM-D
GPTQ/AWQ

Job description

JPMorganChase is seeking a Principal Software Engineer to own LLM inference performance and optimization for scalable AI platforms. You will benchmark, quantify improvements, and drive decisions with leaders to improve speed, cost efficiency, and production readiness across models.

You will collaborate with EKS and disaggregated serving teams, lead quantization experiments, and shape guardrails for secure and auditable AI workflows.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Principal LLM Inference Engineer — AI Platform
Principal LLM Inference Engineer — AI Platform

JPMorganChase • Jersey City (NJ)

On-site
USD 210,000 - 320,000
Senior Principal LLM Engineer — AI Platform & Scale
Senior Principal LLM Engineer — AI Platform & Scale

JPMorgan Chase & Co. • Palo Alto (CA)

On-site
USD 233,000 - 325,000
Principal AI Engineer - Lead LLM Platform & Innovation
Principal AI Engineer - Lead LLM Platform & Innovation

Next Frontier Capital • Palo Alto (CA)

On-site
USD 210,000 - 290,000
Senior Principal AI/ML Engineer — LLM & GNN Architect
Senior Principal AI/ML Engineer — LLM & GNN Architect

Fairygodboss • Palo Alto (CA)

On-site
USD 180,000 - 240,000
GenAI Platform Architect for High-Scale LLM Inference
GenAI Platform Architect for High-Scale LLM Inference

Socket.dev • New Jersey

On-site
USD 180,000 - 240,000
Lead AI Engineer for Enterprise LLM Platform
Lead AI Engineer for Enterprise LLM Platform

Next Frontier Capital • Palo Alto (CA)

On-site
USD 180,000 - 260,000
Senior AI Platform Engineer - Enterprise LLMs
Senior AI Platform Engineer - Enterprise LLMs

Next Frontier Capital • Jersey City (NJ)

On-site
USD 170,000 - 210,000
Principal Software Engineer - LLM Optimization
Principal Software Engineer - LLM Optimization

JPMorgan Chase & Co. • Jersey City (NJ)

On-site
USD 180,000 - 240,000
Senior ML Engineer - LLMs & Production AI for Finance
Senior ML Engineer - LLMs & Production AI for Finance

JPMorgan Chase & Co. • Palo Alto (CA)

On-site
USD 180,000 - 260,000
Lead AI/ML Engineer — LLMs & Production ML on Cloud
Lead AI/ML Engineer — LLMs & Production ML on Cloud

JPMorgan Chase & Co. • Jersey City (NJ)

On-site
USD 120,000 - 180,000