Senior Lead Software Engineer - LLM Ops Platform

JP Morgan Chase

Glasgow

On-site

GBP 62,000 - 102,000

Full time

3 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

JP Morgan Chase's AI and Machine Learning Platform team is hiring to design, build, and operate secure, production-grade AI infrastructure. You will develop backend services, host large language models, and ensure reliability and cost efficiency in cloud and on‑prem environments.

This role emphasizes observability, CI/CD, and scalable GPU/LLM workflows. The position offers full-time work in the UK, with ownership of reliability, performance, and security across the platform, and opportunities to

Qualifications

  • Hands-on experience with system design, development, testing, and production stability.
  • Advanced proficiency in Python for building production-grade services.
  • Experience with automation, CI/CD, and IaC tooling.
  • Knowledge of cloud infrastructure and container orchestration.
  • Strong site reliability engineering practices: incident management, runbooks, reliability patterns.
  • Practical knowledge of observability across metrics, logs, and traces.

Responsibilities

  • Design, develop, troubleshoot, and deliver secure, high-quality production software and services for AI infrastructure.
  • Build backend services and APIs that enable reliable operation of AI infrastructure in production environments.
  • Operate and scale large language model serving infrastructure, including model hosting and routing.
  • Deploy, host, and lifecycle-manage LLMs on cloud and on-premises using reproducible IaC and CD pipelines.
  • Implement observability with dashboards and alerting for LLM and GPU workloads.
  • Lead reliability engineering for LLM endpoints, including capacity planning and incident response.

Skills

System design
Python
Automation
Cloud platforms
IaC
SRE
Observability
Security
AI tooling
Incident management
Model serving
Team collaboration

Tools

Kubernetes
GPU clusters
CI/CD tooling

Job description

Salary: £62,000 - 102,000 per year
Requirements
  • We have hands-on experience with system design, application development, testing, and operational stability in production environments.
  • We have advanced proficiency in Python for building production-grade services and tooling.
  • We have proficiency with automation and continuous delivery methods.
  • We have hands-on experience with cloud infrastructure platforms and infrastructure-as-code tooling for delivery and lifecycle management.
  • We have a strong understanding of site reliability engineering practices, including incident management, root-cause analysis, runbooks, and reliability patterns.
  • We have practical knowledge of observability and instrumentation across metrics, logs, and traces.
  • We have hands-on experience with Kubernetes and container-based orchestration platforms, including managed cloud variants.
  • We have experience hosting and serving large language models on cloud-based infrastructure and local GPU environments.
  • We have knowledge of large language model reliability and risk considerations, including latency and throughput trade-offs, model versioning, prompt and response logging, and safe rollout patterns.
  • We have hands-on experience using enterprise-authorized AI-assisted software development tools within the work environment, with demonstrated ability to critically evaluate, validate, and refine AI-generated outputs for correctness, performance, and security.
  • We understand responsible AI use in engineering workflows, including data sensitivity considerations, secure handling of inputs and outputs, and adherence to resiliency and security expectations.
  • We have the ability to guide peers on safe and effective usage within team practices.
Responsibilities
  • We design, develop, troubleshoot, and deliver secure, high-quality production software and services for AI infrastructure.
  • We build backend services and APIs that enable reliable operation of AI infrastructure in production environments.
  • We operate and scale large language model serving infrastructure, including model hosting, request routing, continuous batching, and cache optimization.
  • We deploy, host, and lifecycle-manage open-source and proprietary large language models on cloud-based container orchestration platforms and on-premises GPU clusters using reproducible infrastructure as code and continuous delivery pipelines.
  • We implement observability across logs, metrics, and traces with dashboards and actionable alerting for large language model and GPU workloads.
  • We tune GPU and accelerator capacity, autoscaling, and cost efficiency for large language model inference workloads using performance optimization techniques such as quantization, parallelism, and speculative decoding.
  • We lead reliability engineering for large language model endpoints through capacity planning, load and soak testing, safe rollouts, failover, and incident response for outages and model-quality regressions.
  • We participate in on-call rotations, lead incident triage and mitigation, and produce clear post-incident root-cause analyses and follow-up actions.
  • We identify recurring operational issues and automate remediation to improve platform stability and developer experience.
  • We build and maintain multi-agent systems with strong orchestration, including planning, coordination, tool-calling, state and memory management, and workflow control where appropriate.
  • We drive team adoption of enterprise-authorized AI-assisted engineering practices to improve code quality, delivery speed, and operational outcomes, while establishing consistent validation standards and promoting reuse of effective patterns across the team.
Technologies
  • AI
  • Backend
  • Cloud
  • Incident Management
  • Support
  • Kubernetes
  • Machine Learning
  • Model Serving
  • Python
  • Security
  • AI Agents
  • LLM
  • Marketing
  • vLLM
More

We are JPMorganChase, a global leader in financial services providing strategic advice and products to prominent corporations, governments, wealthy individuals, and institutional investors. Our AI and Machine Learning Platform team is focused on building and scaling AI infrastructure that modernizes traditional infrastructure management and site reliability engineering through applied AI. We offer a full-time role with meaningful ownership of reliability, performance, and cost-efficiency for large language model inference platforms, along with deep hands-on work in cloud, Kubernetes, observability, and production AI systems. We value diversity and inclusion, support equal opportunity, and provide reasonable accommodations where needed.

last updated 36 week of 2026

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Lead Software Engineer - LLM Ops Platform Reliability
Senior Lead Software Engineer - LLM Ops Platform Reliability

JPMorgan Chase & Co. • Auchentibber

On-site
GBP 90,000 - 150,000
Senior Lead Software Engineer - LLM Ops Platform Reliability
Senior Lead Software Engineer - LLM Ops Platform Reliability

慨正橡扯 • Glasgow

On-site
GBP 90,000 - 130,000
Senior Lead Software Engineer - LLM Ops Platform Reliability
Senior Lead Software Engineer - LLM Ops Platform Reliability

J.P. MORGAN • Bournemouth

On-site
GBP 110,000 - 160,000
Senior Lead Software Engineer - LLM Ops Platform Reliability
Senior Lead Software Engineer - LLM Ops Platform Reliability

J.P. MORGAN • Greater London

On-site
GBP 140,000 - 210,000
Senior Lead Software Engineer - LLM Ops Platform Reliability
Senior Lead Software Engineer - LLM Ops Platform Reliability

JPMorganChase • Glasgow

On-site
GBP 90,000 - 130,000
Senior Lead Software Engineer - LLM Ops Platform Reliability
Senior Lead Software Engineer - LLM Ops Platform Reliability

Jp Morgan Chase • Glasgow

On-site
GBP 110,000 - 150,000
Sr Lead AI Platform Engineer
Sr Lead AI Platform Engineer

JP Morgan Chase • Glasgow

On-site
GBP 62,000 - 102,000
Lead Site Reliability Engineer - Operations Excellence
Lead Site Reliability Engineer - Operations Excellence

JPMorganChase • Glasgow

On-site
GBP 90,000 - 130,000
Lead Software Engineer - Glasgow
Lead Software Engineer - Glasgow

Hackajob Ltd • Glasgow

On-site
GBP 90,000 - 110,000
Lead Software Engineer - Python / Go & AI/ML
Lead Software Engineer - Python / Go & AI/ML

JP Morgan Chase • Glasgow

On-site
GBP 62,000 - 102,000