Data Center Reliability & Automation Engineer

Meta

New Albany (OH)

On-site

USD 111,010 - 158,995

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Meta is seeking a Data Center Production Operations Engineer to support the reliability, efficiency, and scalability of our global data center infrastructure. You will manage day-to-day operational health of server fleets, work across hardware lifecycle management and process improvement, and ensure production environments meet the demands of billions of users.

You will participate in 24/7 on-call rotation and collaborate with hardware, network, and capacity planning teams to optimize

Qualifications

  • 6+ years of data center operations experience.
  • Hands-on server hardware troubleshooting and failure analysis.
  • Experience with monitoring/observability platforms to track fleet health and incidents.
  • Experience creating or improving operational runbooks and automation scripts.
  • Ability to travel up to 15%.

Responsibilities

  • Manage and maintain large-scale server fleets across data center environments, including hardware triage, failure analysis, and coordinating repair and replacement workflows.
  • Monitor production systems health using observability tooling and telemetry data to proactively identify and resolve infrastructure anomalies before they impact service availability
  • Develop and refine operational runbooks, escalation procedures, and incident response playbooks specific to data center server environments
  • Collaborate with hardware engineering, network operations, and capacity planning teams to support server deployment, decommissioning, and lifecycle transitions
  • Analyze failure trends and operational data to identify systemic issues in server hardware or firmware, and drive root cause analysis and corrective action
  • Contribute to automation initiatives that reduce manual toil in server provisioning, health checks, and fleet management workflows, including leveraging AI-integrated tooling
  • Partner with cross-functional teams to evaluate and implement process improvements that increase operational efficiency and reduce mean time to resolution for production incidents
  • Communicate infrastructure status, incident timelines, and risk assessments to engineering and operations stakeholders through clear written and verbal updates
  • Support capacity readiness activities by validating server acceptance criteria and coordinating with data center technicians during hardware bring-up and commissioning
  • Identify gaps in monitoring coverage or operational tooling and propose solutions that improve fleet visibility and production reliability
  • Participate in 24/7 on-call rotation
  • Ability to travel up to 15% of the time

Skills

Data center operations
Server hardware troubleshooting
Observability platforms
Automation scripting
Incident response
Cross-functional collaboration

Tools

Python
Bash
IPMI
Redfish
BIOS configuration

Job description

Meta is seeking a Data Center Production Operations Engineer to support the reliability, efficiency, and scalability of our global data center infrastructure. You will manage day-to-day operational health of server fleets, work across hardware lifecycle management and process improvement, and ensure production environments meet the demands of billions of users.

You will participate in 24/7 on-call rotation and collaborate with hardware, network, and capacity planning teams to optimize

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Center Production & Automation Engineer
Data Center Production & Automation Engineer

Meta • United States

Remote
USD 120,000 - 160,000
Data Center Reliability & Automation Engineer
Data Center Reliability & Automation Engineer

Meta • Huntsville (AL)

On-site
USD 84,000 - 130,000
Data Center Reliability & Operations Engineer
Data Center Reliability & Operations Engineer

Meta • Georgia

On-site
USD 88,000 - 125,000
Data Center Production Ops Engineer | Linux & Hardware
Data Center Production Ops Engineer | Linux & Hardware

Meta • Oregon (WI)

On-site
USD 71,000 - 103,000
Bonus
Equity
Benefits
Data Center Reliability & Automation Engineer
Data Center Reliability & Automation Engineer

Meta • Bowling Green (OH)

On-site
USD 83,000 - 130,000
Equity
Benefits
Bonus
Data Center Reliability Architect
Data Center Reliability Architect

Meta • Augusta (ME)

On-site
USD 143,000 - 198,000
Bonus
Equity
Data Center Reliability Engineer - Asset & Operations
Data Center Reliability Engineer - Asset & Operations

Meta • Atlanta (GA)

On-site
USD 143,000 - 198,000
Data Center Reliability Engineer: Asset & Maintenance Leader
Data Center Reliability Engineer: Asset & Maintenance Leader

Meta • Olympia (WA)

On-site
USD 143,000 - 198,000
Global Data Center Network Engineer & Automation
Global Data Center Network Engineer & Automation

Meta • Jeffersonville (IN)

On-site
USD 120,000 - 180,000
Data Center Reliability & Asset Strategy Lead
Data Center Reliability & Asset Strategy Lead

Meta • United States

On-site
USD 143,000 - 198,000