Data Center Reliability & Automation Engineer

Meta

Newark, New Albany (CA, OH)

On-site

USD 111,000 - 159,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Meta is seeking a Data Center Production Operations Engineer to ensure reliability and scalability of our global data center infrastructure. You will manage large-scale server fleets, monitor fleet health, and drive automation and incident response improvements.

The role requires on-call availability and up to 15% travel. You will collaborate with hardware, network, and capacity teams to deploy, decommission, and lifecycle manage servers, while refining runbooks and tooling to reduce mean time

Qualifications

  • 6+ years of experience in data center operations, site operations, or production infrastructure engineering supporting large-scale server environments.
  • 6+ years of experience with server hardware components including CPUs, memory, storage, and network interface cards, including hands-on troubleshooting and failure diagnosis.
  • Experience using systems monitoring and observability platforms to track fleet health, identify anomalies, and drive incident resolution in production data center environments.
  • Experience developing or improving operational processes, runbooks, or automation scripts to support server fleet management at scale.
  • Experience collaborating with hardware engineering, network, and capacity teams to coordinate infrastructure deployments and lifecycle activities.

Responsibilities

  • Manage and maintain large-scale server fleets across data center environments, including hardware triage, failure analysis, and coordinating repair and replacement workflows.
  • Monitor production systems health using observability tooling and telemetry data to proactively identify and resolve infrastructure anomalies before they impact service availability.
  • Develop and refine operational runbooks, escalation procedures, and incident response playbooks specific to data center server environments.
  • Collaborate with hardware engineering, network operations, and capacity planning teams to support server deployment, decommissioning, and lifecycle transitions.
  • Analyze failure trends and operational data to identify systemic issues in server hardware or firmware, and drive root cause analysis and corrective action.
  • Contribute to automation initiatives that reduce manual toil in server provisioning, health checks, and fleet management workflows, including leveraging AI-integrated tooling.
  • Partner with cross-functional teams to evaluate and implement process improvements that increase operational efficiency and reduce mean time to resolution for production incidents.
  • Communicate infrastructure status, incident timelines, and risk assessments to engineering and operations stakeholders through clear written and verbal updates.
  • Support capacity readiness activities by validating server acceptance criteria and coordinating with data center technicians during hardware bring-up and commissioning.
  • Identify gaps in monitoring coverage or operational tooling and propose solutions that improve fleet visibility and production reliability.
  • Participate in 24/7 on-call rotation.
  • Ability to travel up to 15% of the time.

Skills

Data center operations
Monitoring/observability
Automation scripting

Tools

IPMI/Redfish
Python/Bash scripting
Observability tooling

Job description

Meta is seeking a Data Center Production Operations Engineer to ensure reliability and scalability of our global data center infrastructure. You will manage large-scale server fleets, monitor fleet health, and drive automation and incident response improvements.

The role requires on-call availability and up to 15% travel. You will collaborate with hardware, network, and capacity teams to deploy, decommission, and lifecycle manage servers, while refining runbooks and tooling to reduce mean time

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Center Reliability & Automation Engineer
Data Center Reliability & Automation Engineer

Meta • Huntsville (AL)

On-site
USD 84,000 - 130,000
Data Center Reliability & Operations Engineer
Data Center Reliability & Operations Engineer

Meta • Georgia

On-site
USD 88,000 - 125,000
Data Center Production Ops Engineer | Linux & Hardware
Data Center Production Ops Engineer | Linux & Hardware

Meta • Oregon (WI)

On-site
USD 71,000 - 103,000
Bonus
Equity
Benefits
Data Center Reliability & Automation Engineer
Data Center Reliability & Automation Engineer

Meta • Bowling Green (OH)

On-site
USD 83,000 - 130,000
Equity
Benefits
Bonus
Data Center Production Operations Engineer
Data Center Production Operations Engineer

Meta • Hillsboro (OR)

Hybrid
USD 95,000 - 125,000
Data Center Reliability & Asset Strategy Lead
Data Center Reliability & Asset Strategy Lead

Meta • United States

On-site
USD 143,000 - 198,000
Data Center Network & Automation Engineer
Data Center Network & Automation Engineer

Meta • Menlo Park (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Hyperscale Production Systems Engineer
Hyperscale Production Systems Engineer

Meta • Nebraska

On-site
USD 144,000 - 204,000
Data Center Production Ops Engineer (Linux & Hardware)
Data Center Production Ops Engineer (Linux & Hardware)

Meta • Hillsboro (OR)

Hybrid
USD 95,000 - 125,000
Data Center Production Operations Engineer
Data Center Production Operations Engineer

Meta • Newark (CA), New Albany (OH)

On-site
USD 111,000 - 159,000