Senior Platform Reliability Engineer for AI Infrastructure

J.P. MORGAN

Bournemouth

On-site

GBP 110,000 - 160,000

Full time

13 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

J.P. Morgan is hiring a Senior Lead Software Engineer to shape AI infrastructure and scale AI platforms in production.

You will own reliability, performance, and cost-efficiency of large language model inference stacks, work with Kubernetes-based deployments, and drive observability and incident response across distributed systems. You will collaborate with engineering to deliver secure software, perform capacity planning, and lead on-call rotations, while continuously improving platform

Qualifications

  • Hands-on experience with system design, application development, testing, and operational stability in production environments.
  • Advanced proficiency in Python for building production-grade services and tooling.
  • Proficiency with automation and continuous delivery methods.
  • Hands-on experience with cloud infrastructure platforms and infrastructure-as-code tooling for delivery and lifecycle management.
  • Strong understanding of site reliability engineering practices, including incident management, root-cause analysis, runbooks, and reliability patterns.
  • Practical knowledge of observability and instrumentation across metrics, logs, and traces.
  • Hands-on experience with Kubernetes and container-based orchestration platforms, including managed cloud variants.
  • Experience hosting and serving large language models on cloud-based infrastructure and local GPU environments.
  • Knowledge of large language model reliability and risk considerations, including latency and throughput trade-offs, model versioning, prompt and response logging, and safe rollout patterns.
  • Hands-on experience using enterprise-authored AI-assisted software development tools within the work environment with demonstrated ability to critically evaluate, validate, and refine AI-generated outputs for correctness, performance, and security.
  • Understanding of responsible AI use in engineering workflows, including data sensitivity considerations, secure handling of inputs/outputs, and adherence to resiliency and security expectations; ability to guide peers on safe and effective usage within team practices.

Responsibilities

  • Design, develop, troubleshoot, and deliver secure, high‑quality production software and services for AI infrastructure
  • Build backend services and APIs that enable reliable operation of AI infrastructure in production environments
  • Operate and scale large language model serving infrastructure, including model hosting, request routing, continuous batching, and cache optimization
  • Deploy, host, and lifecycle‑manage open‑source and proprietary large language models on cloud‑based container orchestration platforms and on‑premises GPU clusters using reproducible infrastructure as code and continuous delivery pipelines
  • Implement observability across logs, metrics, and traces with dashboards and actionable alerting for large language model and GPU workloads
  • Tune GPU and accelerator capacity, autoscaling, and cost efficiency for large language model inference workloads using performance optimization techniques such as quantization, parallelism, and speculative decoding
  • Lead reliability engineering for large language model endpoints through capacity planning, load and soak testing, safe rollouts, failover, and incident response for outages and model‑quality regressions
  • Participate in on‑call rotations, lead incident triage and mitigation, and produce clear post‑incident root‑cause analyses and follow‑up actions
  • Identify recurring operational issues and automate remediation to improve platform stability and developer experience
  • Build and maintain multi‑agent systems with strong orchestration, including planning, coordination, tool‑calling, state and memory management, and workflow control where appropriate
  • Drives team adoption of enterprise‑authorized AI‑assisted engineering practices within the work environment to improve code quality, delivery speed, and operational outcomes while establishing consistent validation standards and promoting reuse of patterns across the team.

Skills

Python
Kubernetes
Cloud infra
Observability
Site reliability
Distributed systems
GPU/ML infra

Job description

J.P. Morgan is hiring a Senior Lead Software Engineer to shape AI infrastructure and scale AI platforms in production.

You will own reliability, performance, and cost-efficiency of large language model inference stacks, work with Kubernetes-based deployments, and drive observability and incident response across distributed systems. You will collaborate with engineering to deliver secure software, perform capacity planning, and lead on-call rotations, while continuously improving platform

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Lead, AI Infra Reliability Engineer
Senior Lead, AI Infra Reliability Engineer

JPMorganChase • Glasgow

On-site
GBP 90,000 - 130,000
Senior Lead SRE: AI Infra & LLM Serving at Scale
Senior Lead SRE: AI Infra & LLM Serving at Scale

J.P. MORGAN • Greater London

On-site
GBP 140,000 - 210,000
Senior AI/ML Platform Reliability Engineer
Senior AI/ML Platform Reliability Engineer

J.P. MORGAN • Scotland

On-site
GBP 70,000 - 100,000
Senior Platform Engineer for AI/ML Production & Reliability
Senior Platform Engineer for AI/ML Production & Reliability

Fairygodboss • Glasgow

On-site
GBP 180,000 - 240,000
Senior AI/ML Platform Engineer
Senior AI/ML Platform Engineer

JPMorganChase • Greater London

On-site
GBP 100,000 - 140,000
Senior AI Platform Engineer: Scale, Reliability & Ops
Senior AI Platform Engineer: Scale, Reliability & Ops

JPMorganChase • Glasgow

On-site
GBP 180,000 - 240,000
Senior Lead AI Infrastructure Reliability & LLM Serving
Senior Lead AI Infrastructure Reliability & LLM Serving

JPMorgan Chase & Co. • Glasgow

On-site
GBP 90,000 - 120,000
Senior Lead Software Engineer - LLM Ops Platform Reliability
Senior Lead Software Engineer - LLM Ops Platform Reliability

JPMorgan Chase & Co. • Glasgow

On-site
GBP 90,000 - 120,000
Senior AI Platform Engineer - Gen AI & Cloud Security
Senior AI Platform Engineer - Gen AI & Cloud Security

J.P. MORGAN • Greater London

On-site
GBP 90,000 - 150,000
Senior AI/ML Platform Reliability Engineer
Senior AI/ML Platform Reliability Engineer

J.P. MORGAN • Bournemouth

On-site
GBP 70,000 - 110,000