Infrastructure / DevOps Lead

DNEG

Greater London

On-site

GBP 150,000 - 190,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Brahma AI is building state‑of‑the‑art generative AI platforms, powering AI models and GPU‑heavy workloads across multi‑cloud environments. We seek a Lead Infrastructure/DevOps Engineer to drive a 7‑person team and own the core platform infrastructure for scalable training and deployment pipelines.

You will guide architecture, implement IaC, and optimize cloud costs while ensuring security and reliability for enterprise AI research and production workloads.

Qualifications

  • Leadership experience leading 5+ infra/DevOps engineers.
  • Hands-on AI/GPU infra for compute-heavy workloads.
  • Proficiency with Kubernetes, Docker, Terraform/OpenTofu, and multi‑cloud (GCP).
  • Designing/maintaining high‑performance storage architectures.
  • Cloud cost governance, FinOps, capacity planning, vendor management.
  • Strong communication bridging business needs and technical infra.

Responsibilities

  • Lead and mentor a team of DevOps and Infrastructure engineers.
  • Drive agile delivery, sprint planning, and backlog prioritisation.
  • Set Reliability Engineering, IaC, CI, and incident post‑mortems standards.
  • Act as technical bridge between infra, ML researchers, and AI apps.
  • Oversee GPU clusters and multi‑cloud infra for model training and inference.
  • Manage FinOps, cloud/hardware costs and vendor terms; optimise GPU use.
  • Standardise deployments with Terraform/OpenTofu and Kubernetes.
  • Collaborate on security to meet ISO 27001, SOC 2 and related standards.

Skills

Leadership
AI/GPU infra
Multi-Cloud
High‑Performance Storage
FinOps / Cloud Cost
Communication

Tools

Kubernetes
Docker
Terraform/OpenTofu
GCP

Job description

Brahma AI operates at the intersection of enterprise Media Asset Management (MAM) and cutting‑edge generative media. We build and scale industry‑leading generative AI models, including hyper‑realistic digital humans (ATMAN) and multilingual voice synthesis (VAANI), for world‑class enterprise clients in entertainment, sports, healthcare, and retail.

Role Overview

We are looking for a Lead Infrastructure / DevOps Engineer to lead our core AI Platform & Infrastructure team. Reporting directly to the VP of Engineering, you will guide a team of 7 engineers responsible for powering our high‑performance GPU infrastructure, multi‑cloud setup (with GCP as our primary provider), ML model training pipelines, and containerised orchestration environments. This role balances technical leadership, team management, cloud resource optimisation, and high‑level architectural oversight. You will ensure our research and engineering teams have the fast, scalable, and reliable compute environments necessary to train and deploy state‑of‑the‑art AI models.

Key Responsibilities
  • People, Team & Process Leadership (50%)
  • Team Management: Lead, mentor, and grow a team of 7 DevOps and Infrastructure engineers through regular 1:1s, performance reviews, and career pathing.
  • Sprint & Operational Delivery: Drive agile delivery, sprint planning, and backlog prioritisation to align infrastructure deliverables with AI research and product roadmaps.
  • Engineering Standards: Establish best practices for Reliability Engineering, Infrastructure‑as‑Code (IaC), continuous integration, and incident post‑mortems.
  • Cross‑Functional Alignment: Act as the primary technical bridge between infrastructure, ML researchers, AI application developers, and the VP of Engineering.
  • Resource Management & FinOps (25%)
  • GPU & Multi‑Cloud Management: Manage high‑density GPU clusters across a multi‑cloud ecosystem (primarily GCP) optimised for large custom AI model training and real‑time inference workflows.
  • FinOps & Cost Control: Oversee infrastructure consumption, track cloud/hardware costs, negotiate vendor terms, and optimise GPU utilisation to maintain cost efficiency.
  • High‑Performance Storage: Oversee high‑throughput storage and caching solutions engineered for ultra‑fast data retrieval and low‑latency access.
  • Architecture, Engineering & Compliance (25%)
  • Technical Escalation & Hands‑On Oversight: Serve as the senior technical escalation point for complex infrastructure incidents and architecture decisions.
  • Automation & IaC: Standardise platform deployments using Infrastructure as Code (e.g., Terraform/OpenTofu) and modern container orchestration (Kubernetes).
  • Security & Compliance Collaboration: Partner with security stakeholders to ensure our AI training environments meet industry security standards (e.g., MPA Best Practices, ISO 27001, SOC 2).
Must Haves
  • Leadership Experience: Proven track record leading or managing a team of 5+ infrastructure, platform, or DevOps engineers.
  • AI/GPU Infrastructure: Hands‑on experience architecting and managing GPU‑intensive workloads (NVIDIA clusters, cloud AI accelerators) for compute‑heavy applications.
  • Multi‑Cloud & Orchestration: Expertise with Kubernetes, Docker, Terraform (or OpenTofu), and multi‑cloud environments (with strong hands‑on GCP experience).
  • High‑Performance Storage: Demonstrated experience designing, optimising, and maintaining high‑performance storage architectures and caching layers for demanding compute workloads.
  • Cloud & Resource Management: Strong experience with cloud cost governance (FinOps), capacity planning, and vendor interaction.
  • Communication: Exceptional stakeholder management skills with the ability to bridge business requirements and deep technical infrastructure details.
Nice to Have
  • Experience managing physical data centres, co‑location facilities, or hybrid infrastructure environments.
  • Working knowledge of ML orchestration frameworks (e.g., Ray, Slurm, Kubeflow).
  • Background in media pipelines, VFX tooling, or media compliance standards (MPA, ISO 27001).
  • Prior experience working in a hybrid startup/scale‑up environment.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Backend Engineer
Senior Backend Engineer

DNEG • Greater London

On-site
GBP 90,000 - 120,000
Senior AI Infrastructure & DevOps Lead
Senior AI Infrastructure & DevOps Lead

DNEG • Greater London

On-site
GBP 150,000 - 190,000
Infrastructure Lead
Infrastructure Lead

adm-Indicia • Greater London

On-site
GBP 110,000 - 140,000
Infrastructure Lead
Infrastructure Lead

adm Indicia • Greater London

On-site
GBP 90,000 - 130,000
Advisory AI Infrastructure Engineer
Advisory AI Infrastructure Engineer

Lenovo • City of Edinburgh

Hybrid
GBP 70,000 - 90,000
Opportunities for career advancement
Diverse training programs
Performance-based rewards
+3
AI Infrastructure Architect
AI Infrastructure Architect

Accenture UK & Ireland • Greater London

On-site
GBP 120,000 - 170,000
Infrastructure Lead
Infrastructure Lead

Indicia Ltd • Greater London

Hybrid
GBP 110,000 - 150,000
Senior Platform Engineer (Product Initiatives) - Systems Integrator
Senior Platform Engineer (Product Initiatives) - Systems Integrator

Hamilton Barnes Associates Limited • United Kingdom

Remote
GBP 120,000 - 190,000
High-Upside Equity
Flexible remote setup
Work-Life Balance
+1
AI DevOps Engineer
AI DevOps Engineer

Bluecrestcapitalmanagement • Greater London

On-site
GBP 90,000 - 150,000
Platform Lead - ML Ops
Platform Lead - ML Ops

Anaplan Inc • Greater London

Hybrid
GBP 120,000 - 180,000