Data Center Production Operations Engineer

Meta

Clonee

On-site

EUR 70,000 - 100,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Meta seeks a Data Center Production Operations Engineer to ensure reliability, efficiency, and scalability of its global data center infrastructure. You will manage the day-to-day operational health of server fleets and production systems across environments.

You will partner with hardware engineering, capacity planning, and infrastructure teams to keep data centers at peak performance, enabling Meta's products and services used worldwide.

Qualifications

  • 2+ years of experience in data center operations, production operations, or systems administration in a large-scale infrastructure environment.
  • Experience troubleshooting server hardware components including CPUs, memory, storage, and networking hardware in a production setting.
  • Experience developing or improving operational processes, runbooks, or standard operating procedures for data center or infrastructure teams.
  • Experience analyzing operational data or failure metrics to identify trends and drive reliability improvements.
  • Experience collaborating with cross-functional engineering teams to resolve production incidents and implement systemic fixes.

Responsibilities

  • Monitor and maintain the operational health of large-scale server fleets and production infrastructure across data center environments
  • Diagnose and resolve hardware and systems failures, coordinating with engineering teams to drive root cause analysis and implement corrective actions
  • Execute and refine server deployment, decommissioning, and lifecycle management processes to support capacity and reliability goals
  • Develop and maintain operational runbooks, escalation procedures, and documentation to standardize production operations workflows
  • Collaborate with hardware engineering and capacity planning teams to identify systemic issues and propose infrastructure improvements
  • Track and analyze operational metrics and failure trends to surface insights that improve fleet reliability and reduce mean time to resolution
  • Support the qualification and rollout of new server hardware generations by validating operational readiness and identifying deployment risks
  • Partner with cross-functional teams including network engineering, facilities, and software infrastructure to resolve complex production incidents
  • Identify opportunities to automate repetitive operational tasks and contribute to tooling improvements that increase operational efficiency
  • Provide technical guidance to peers on production operations best practices, hardware troubleshooting methodologies, and process standards

Skills

Data center ops
Server hardware troubleshooting
Operational processes
Failure analysis
Cross-functional collaboration

Job description

Summary

Meta is seeking a Data Center Production Operations Engineer to support the reliability, efficiency, and scalability of our global data center infrastructure. In this role, you will be responsible for the day-to-day operational health of server fleets and production systems, working at the intersection of hardware lifecycle management, systems troubleshooting, and operational process improvement. You will partner closely with hardware engineering, capacity planning, and infrastructure teams to ensure Meta's data centers operate at peak performance, directly enabling the products and services that connect billions of people worldwide.

Required Skills
Data Center Production Operations Engineer Responsibilities
  1. Monitor and maintain the operational health of large-scale server fleets and production infrastructure across data center environments
  2. Diagnose and resolve hardware and systems failures, coordinating with engineering teams to drive root cause analysis and implement corrective actions
  3. Execute and refine server deployment, decommissioning, and lifecycle management processes to support capacity and reliability goals
  4. Develop and maintain operational runbooks, escalation procedures, and documentation to standardize production operations workflows
  5. Collaborate with hardware engineering and capacity planning teams to identify systemic issues and propose infrastructure improvements
  6. Track and analyze operational metrics and failure trends to surface insights that improve fleet reliability and reduce mean time to resolution
  7. Support the qualification and rollout of new server hardware generations by validating operational readiness and identifying deployment risks
  8. Partner with cross-functional teams including network engineering, facilities, and software infrastructure to resolve complex production incidents
  9. Identify opportunities to automate repetitive operational tasks and contribute to tooling improvements that increase operational efficiency
  10. Provide technical guidance to peers on production operations best practices, hardware troubleshooting methodologies, and process standards
Minimum Qualifications
  1. 2+ years of experience in data center operations, production operations, or systems administration in a large-scale infrastructure environment
  2. Experience troubleshooting server hardware components including CPUs, memory, storage, and networking hardware in a production setting
  3. Experience developing or improving operational processes, runbooks, or standard operating procedures for data center or infrastructure teams
  4. Experience analyzing operational data or failure metrics to identify trends and drive reliability improvements
  5. Experience collaborating with cross-functional engineering teams to resolve production incidents and implement systemic fixes
Preferred Qualifications
  1. Experience supporting hardware qualification or new server platform bring-up in a data center production environment
  2. Experience with fleet management tooling, asset tracking systems, or infrastructure monitoring platforms at scale
  3. Familiarity with scripting languages such as Python or Bash for automating operational workflows and data analysis tasks
  4. Background in capacity planning, hardware lifecycle management, or server deployment operations for hyperscale data centers

Industry: Internet

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Data Center Production & Reliability Engineer
Data Center Production & Reliability Engineer

Meta • Clonee

On-site
EUR 70,000 - 100,000
Data Center Production Operations Engineer
Data Center Production Operations Engineer

Meta Careers • Clonee

On-site
EUR 65,000 - 95,000
Production Systems Engineer: Scale & Automate Global Fleet
Production Systems Engineer: Scale & Automate Global Fleet

Meta • Dublin

On-site
EUR 90,000 - 150,000
Data Center Operations Engineer: Reliability at Scale
Data Center Operations Engineer: Reliability at Scale

Meta Careers • Clonee

On-site
EUR 65,000 - 95,000
Production Systems Engineer
Production Systems Engineer

Meta • Dublin

On-site
EUR 90,000 - 150,000
Production Systems Engineer: Scale & Automate the Fleet
Production Systems Engineer: Scale & Automate the Fleet

Meta Careers • Dublin

On-site
EUR 90,000 - 130,000
Production Systems Engineer
Production Systems Engineer

Meta Careers • Dublin

On-site
EUR 90,000 - 130,000
Data Center Technical Operations Engineer
Data Center Technical Operations Engineer

GemPool Recruitment • Blanchardstown

On-site
EUR 60,000 - 82,000
Data Center Engineer: Deploy, Maintain & Optimize Infra
Data Center Engineer: Deploy, Maintain & Optimize Infra

FRONTIER FORCE TECHNOLOGY PTE LTD • Dublin

On-site
EUR 45,000 - 65,000
Data Center Engineer
Data Center Engineer

FRONTIER FORCE TECHNOLOGY PTE LTD • Dublin

On-site
EUR 45,000 - 65,000