Senior Lead AI Infra Engineer for LLM Ops

Hackajob Ltd

Glasgow

On-site

GBP 90,000 - 110,000

Full time

2 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

hackajob Ltd. in the United Kingdom seeks a Senior Lead Software Engineer for the AI and ML Platform team to design and scale production AI infrastructure. You will lead reliability for LLM endpoints and build backend services for dependable AI operations.

Responsibilities include deploying LLMs on cloud and on-prem GPUs, managing containers via Kubernetes, and implementing observability, CD pipelines, and cost-efficient scaling.

Qualifications

  • Advanced Python for production-grade services
  • Strong SRE practices incl. incident management and runbooks
  • Experience with cloud infra and IaC tooling for lifecycle management
  • Kubernetes and container orchestration in cloud and on-prem
  • Hosting and serving large language models on cloud and local GPUs
  • Observability across metrics, logs, and traces

Responsibilities

  • Design, develop, troubleshoot, and deliver production software for AI infra
  • Build backend services and APIs for reliable AI infrastructure
  • Operate and scale LLM serving infra across cloud and on-prem clusters
  • Lifecycle-manage open-source and proprietary LLMs using IaC and CD pipelines
  • Implement observability dashboards and alerting for GPUs/workloads
  • Lead reliability engineering for LLM endpoints, capacity planning, and incident response
  • Automate remediation to improve platform stability and developer experience
  • Drive enterprise AI practices to improve code quality and delivery speed

Skills

Python
Kubernetes
Cloud
Machine Learning
Reliability Engineering
Observability
AI/ML Platform

Tools

Infrastructure as Code
CI/CD
Monitoring

Job description

hackajob Ltd. in the United Kingdom seeks a Senior Lead Software Engineer for the AI and ML Platform team to design and scale production AI infrastructure. You will lead reliability for LLM endpoints and build backend services for dependable AI operations.

Responsibilities include deploying LLMs on cloud and on-prem GPUs, managing containers via Kubernetes, and implementing observability, CD pipelines, and cost-efficient scaling.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Lead LLM Ops Platform Reliability Engineer
Lead LLM Ops Platform Reliability Engineer

Hackajob Ltd • Glasgow

On-site
GBP 85,000 - 115,000
Senior LLM Serving Platform Engineer
Senior LLM Serving Platform Engineer

Scale AI • Greater London

On-site
GBP 90,000 - 130,000
Senior Lead Software Engineer - AI-Driven Cloud Platform
Senior Lead Software Engineer - AI-Driven Cloud Platform

Hackajob Ltd • City of Westminster

On-site
GBP 90,000 - 110,000
AI Infra Engineer for Scalable LLM Serving
AI Infra Engineer for Scalable LLM Serving

Scale AI, Inc. • Greater London

On-site
GBP 110,000 - 160,000
Lead Software Engineer - LLM Ops Platform Reliability
Lead Software Engineer - LLM Ops Platform Reliability

Hackajob Ltd • Glasgow

On-site
GBP 85,000 - 115,000
ML Engineering Lead: Production AI & RAG for LegalTech
ML Engineering Lead: Production AI & RAG for LegalTech

Hackajob Ltd • Farringdon

On-site
GBP 85,000 - 115,000
Senior AI Infra Engineer – LLM Ops & Reliability
Senior AI Infra Engineer – LLM Ops & Reliability

JPMorgan Chase & Co. • Auchentibber

On-site
GBP 90,000 - 150,000
Senior MLOps Engineer — Lead AI Platform Infra (Hybrid London)
Senior MLOps Engineer — Lead AI Platform Infra (Hybrid London)

Harnham - Data & Analytics Recruitment • Greater London

Hybrid
GBP 75,000 - 85,000
Lead ML Engineer - Scalable AI Platforms
Lead ML Engineer - Scalable AI Platforms

Faculty • Greater London

Hybrid
GBP 110,000 - 140,000
AI Infra Engineer: Scalable LLM Serving & Platform Design
AI Infra Engineer: Scalable LLM Serving & Platform Design

scaleai • Greater London

On-site
GBP 110,000 - 160,000