Production Engineering Manager

Meta

Greater London

Hybrid

GBP 120,000 - 170,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Meta in London seeks a Production Engineering Manager to lead a team responsible for reliability, scalability, and operational excellence of Meta's production infrastructure and services. You will own the full lifecycle of systems, from capacity planning to incident response and automation.

You will drive technical strategy, champion AI-augmented workflows, and partner with software engineering, data science, and product teams to ensure services operate at global scale with high availability and

Qualifications

  • 4+ years in production engineering, site reliability engineering, or systems software engineering.
  • 2+ years managing production or infrastructure engineering teams.
  • Experience driving reliability and scalability for large-scale distributed systems.
  • Experience coding in Python, C++, Go, or Bash and collaborating technically.
  • Ability to set goals, manage roadmaps, and communicate infrastructure strategy.

Responsibilities

  • Manage a team delivering reliability, scalability, and efficiency across interdependent systems.
  • Drive roadmap for infrastructure reliability, capacity planning, and automation.
  • Lead adoption of AI-augmented workflows across the team.
  • Contribute hands-on to code, design reviews, and incident response.
  • Partner with software, data science, and product teams to unblock dependencies.
  • Identify and reduce on-call toil and automate repetitive tasks.
  • Set goals, give feedback, and develop engineers' AI skills.
  • Establish and monitor SLAs, reliability metrics, and engineering efficiency.
  • Communicate health and strategy to leadership and stakeholders.
  • Recruit and retain production engineering talent and minimize single points of failure.

Skills

Reliability engineering
Scalability
Operational efficiency
Team management
AI-augmented workflows
Incident response
Capacity planning
Automation
Code reviews
System design

Tools

Python
C++
Go
Bash

Job description

Summary:

Meta is seeking a Production Engineering Manager to lead a team responsible for the reliability, scalability, and operational excellence of Meta's production infrastructure and services. In this role, you will manage a team of production engineers who own the full lifecycle of systems — from capacity planning and performance optimization to incident response and automation. You will drive technical strategy, champion AI-augmented workflows, and partner closely with software engineering, infrastructure, and product teams to ensure Meta's services operate at global scale with high availability and efficiency.

Required Skills:

Production Engineering Manager Responsibilities:

  1. Manage a team of production engineers delivering on reliability, scalability, and operational efficiency across multiple interdependent production systems
  2. Drive roadmap creation for infrastructure reliability initiatives, capacity planning, and automation efforts, increasing team scope as AI-driven productivity improves throughput
  3. Lead adoption of AI-augmented engineering workflows across the team, sharing learnings and best practices with the broader production engineering organization
  4. Contribute hands-on to technical work including code, system design reviews, and incident response, using AI tooling to expand personal and team reach across disciplines
  5. Partner cross-functionally with software engineering, data science, and product teams to unblock dependencies and ensure smooth execution of infrastructure and reliability projects
  6. Proactively identify and resolve sources of operational toil — including on-call load, alerting gaps, and technical debt — and implement automation to increase team efficiency and scope
  7. Set clear goals and expectations for individual team members, provide timely and actionable feedback, and actively develop engineers' skills including proficiency with AI-augmented workflows
  8. Establish and monitor service-level objectives, reliability metrics, and engineering efficiency indicators to maintain high engineering craft and product quality
  9. Communicate production system health, incident learnings, and infrastructure strategy effectively to engineering leadership and cross-functional stakeholders
  10. Recruit, onboard, and retain production engineering talent, ensuring the team structure minimizes single points of failure and supports sustainable growth
Minimum Qualifications:
  1. 4+ years of experience in production engineering, site reliability engineering, or systems software engineering
  2. 2+ years of experience managing production engineering or infrastructure engineering teams
  3. Experience driving reliability and scalability improvements for large-scale distributed systems, including incident management, capacity planning, and performance optimization
  4. Experience coding and debugging in at least one systems or scripting language (such as Python, C++, Go, or Bash) and contributing technically alongside a team
  5. Experience setting team goals, managing execution against roadmaps, and communicating infrastructure strategy to technical and non-technical stakeholders
Preferred Qualifications:
  1. Demonstrated ongoing AI skill development (e.g., prompt/context engineering, agent orchestration) and staying current with emerging AI technologies
  2. Track record of cross-functional collaboration with software engineering and data science teams to co-own reliability outcomes for consumer-facing or infrastructure services
  3. Demonstrated ability to integrate AI tools to optimize/redesign workflows and drive measurable impact (e.g., efficiency gains, quality improvements)
  4. Experience leading adoption of AI-assisted tooling or automation frameworks within an engineering team to expand operational scope and reduce toil
  5. Experience managing on-call rotations, defining service-level objectives, and implementing observability and alerting improvements at scale
  6. Familiarity with container orchestration, service mesh architectures, or large-scale deployment pipelines in a production environment
  7. Experience adhering to and implementing responsible, ethical AI practices (e.g., risk assessment, bias mitigation, quality and accuracy reviews)
Industry:

Internet

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Production Engineering Manager
Production Engineering Manager

Meta • City of Westminster

On-site
GBP 110,000 - 150,000
Global-Scale Production Engineering Leader
Global-Scale Production Engineering Leader

Meta • Greater London

Hybrid
GBP 120,000 - 170,000
Production Engineering Lead - AI-Driven Reliability
Production Engineering Lead - AI-Driven Reliability

Meta • City of Westminster

On-site
GBP 110,000 - 150,000
Product Manager (Leadership) - Risk
Product Manager (Leadership) - Risk

Meta • City of Westminster

On-site
GBP 120,000 - 180,000
Production Engineering Manager, Rotational Network Engineering (RNE) Program
Production Engineering Manager, Rotational Network Engineering (RNE) Program

Meta • Greater London

On-site
GBP 90,000 - 130,000
Production AI Engineer - Vice President
Production AI Engineer - Vice President

Citigroup Inc. • Greater London

Hybrid
GBP 90,000 - 120,000
27 days annual leave
Discretional annual bonus
Private Medical Care & Life Insurance
+5
Software Engineer, Infrastructure
Software Engineer, Infrastructure

Meta • Greater London

On-site
GBP 90,000 - 150,000
Business Engineer
Business Engineer

Meta • Greater London

Hybrid
GBP 100,000 - 140,000
Senior Production Engineer
Senior Production Engineer

Clear Street • Greater London

On-site
GBP 90,000 - 130,000
Production Manager
Production Manager

Alloyed • Oxford

On-site
GBP 65,000 - 90,000