Data Center Production & Reliability Engineer

Meta

Clonee

On-site

EUR 70,000 - 100,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Meta seeks a Data Center Production Operations Engineer to ensure reliability, efficiency, and scalability of its global data center infrastructure. You will manage the day-to-day operational health of server fleets and production systems across environments.

You will partner with hardware engineering, capacity planning, and infrastructure teams to keep data centers at peak performance, enabling Meta's products and services used worldwide.

Qualifications

  • 2+ years of experience in data center operations, production operations, or systems administration in a large-scale infrastructure environment.
  • Experience troubleshooting server hardware components including CPUs, memory, storage, and networking hardware in a production setting.
  • Experience developing or improving operational processes, runbooks, or standard operating procedures for data center or infrastructure teams.
  • Experience analyzing operational data or failure metrics to identify trends and drive reliability improvements.
  • Experience collaborating with cross-functional engineering teams to resolve production incidents and implement systemic fixes.

Responsibilities

  • Monitor and maintain the operational health of large-scale server fleets and production infrastructure across data center environments
  • Diagnose and resolve hardware and systems failures, coordinating with engineering teams to drive root cause analysis and implement corrective actions
  • Execute and refine server deployment, decommissioning, and lifecycle management processes to support capacity and reliability goals
  • Develop and maintain operational runbooks, escalation procedures, and documentation to standardize production operations workflows
  • Collaborate with hardware engineering and capacity planning teams to identify systemic issues and propose infrastructure improvements
  • Track and analyze operational metrics and failure trends to surface insights that improve fleet reliability and reduce mean time to resolution
  • Support the qualification and rollout of new server hardware generations by validating operational readiness and identifying deployment risks
  • Partner with cross-functional teams including network engineering, facilities, and software infrastructure to resolve complex production incidents
  • Identify opportunities to automate repetitive operational tasks and contribute to tooling improvements that increase operational efficiency
  • Provide technical guidance to peers on production operations best practices, hardware troubleshooting methodologies, and process standards

Skills

Data center ops
Server hardware troubleshooting
Operational processes
Failure analysis
Cross-functional collaboration

Job description

Meta seeks a Data Center Production Operations Engineer to ensure reliability, efficiency, and scalability of its global data center infrastructure. You will manage the day-to-day operational health of server fleets and production systems across environments.

You will partner with hardware engineering, capacity planning, and infrastructure teams to keep data centers at peak performance, enabling Meta's products and services used worldwide.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Data Center Production Operations Engineer
Data Center Production Operations Engineer

Meta • Clonee

On-site
EUR 70,000 - 100,000
Data Center Production Operations Engineer
Data Center Production Operations Engineer

Meta Careers • Clonee

On-site
EUR 65,000 - 95,000
Data Center Operations Engineer: Reliability at Scale
Data Center Operations Engineer: Reliability at Scale

Meta Careers • Clonee

On-site
EUR 65,000 - 95,000
Production Systems Engineer: Scale & Automate Global Fleet
Production Systems Engineer: Scale & Automate Global Fleet

Meta • Dublin

On-site
EUR 90,000 - 150,000
Production Systems Engineer: Scale & Automate the Fleet
Production Systems Engineer: Scale & Automate the Fleet

Meta Careers • Dublin

On-site
EUR 90,000 - 130,000
Production Systems Engineer
Production Systems Engineer

Meta • Dublin

On-site
EUR 90,000 - 150,000
Production Systems Engineer
Production Systems Engineer

Meta Careers • Dublin

On-site
EUR 90,000 - 130,000
Scale-Driven Network Engineer: Data Centers & Automation
Scale-Driven Network Engineer: Data Centers & Automation

Meta • Dublin

On-site
EUR 60,000 - 90,000
Site Reliability & Production Operations Engineer
Site Reliability & Production Operations Engineer

RECRUITERS • Dublin

On-site
EUR 70,000 - 120,000
Senior M&E Reliability Engineer - Data Center AI
Senior M&E Reliability Engineer - Data Center AI

Nscale • Dublin

Hybrid
EUR 80,000 - 110,000
Base + bonus + equity
Dynamic progression
Flexible work environment
+1