Global Production Systems Engineer

Meta

Nebraska

On-site

USD 144,000 - 204,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Meta is seeking an experienced Production Systems Engineer to join the Data Center Operations team. Our data centers and thousands of servers form the foundation for Meta's rapidly scaling infrastructure and services.

The candidate will analyze data, write automation tooling in Python/Bash/C/C++, and drive highly available server repair solutions across a hyperscale fleet, collaborating with globally distributed teams and mentoring other engineers. Travel up to 25% for new site deployments.

Qualifications

  • Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience.
  • 6+ years of experience in production systems engineering, infrastructure engineering, or systems software development for large-scale hardware environments.
  • 6+ years of experience with hardware lifecycle management, fleet automation, or data center operations systems spanning compute, storage, or networking infrastructure.
  • Experience developing systems software or automation tooling in Python, Bash, PHP, C, or C++ for Linux-based production environments at scale.
  • Experience with configuration and maintenance of production systems including web servers, load balancers, relational databases, storage systems, and messaging systems.
  • Experience communicating technical designs and infrastructure decisions through written documentation and cross-functional stakeholder alignment across engineering and operations teams.

Responsibilities

  • Identify and root cause systemic issues across the server fleet and drive resolutions to maximize uptime and utilization by leveraging hardware failure data and diagnostic telemetry.
  • Write, review, and maintain code for diagnostic and automation tooling that supports quality and efficient delivery of production servers at hyperscale.
  • Own and develop diagnostic tooling requirements that enable frontline operations teams to efficiently manage and repair the server fleet.
  • Drive the escalation process for Data Center Operations to identify, root cause, and resolve complex tooling and hardware issues affecting fleet health.
  • Execute operational validation and verification activities for new product integration into the production environment.
  • Collaborate with cross-functional tooling teams to provide an operations-centric perspective on open issues and contribute to their development roadmaps.
  • Perform deep data analysis to prioritize automation opportunities for server repair workflows in a large-scale, heterogeneous hardware environment.
  • Build cross-functional relationships and influence policies and procedures to improve global data center operations consistency and efficiency.
  • Mentor other engineers on evaluating and resolving fleet issues and defining improvements to tools and operational processes.
  • Travel up to 25% to support global data center operations and new site deployments

Skills

Python
Bash
C/C++
Linux
Automation tooling
Data center ops
Documentation

Education

Bachelor's degree

Tools

Web servers
Load balancers
Relational databases
Messaging systems

Job description

Meta is seeking an experienced Production Systems Engineer to join the Data Center Operations team. Our data centers and the tens of thousands of servers installed within them form the foundation upon which Meta's rapidly scaling infrastructure operates and upon which innovative services are delivered. Meta is at the leading edge of the global data center industry in both design and operations. This role requires a forward-thinking systems professional with deep experience leveraging diverse software tools to identify automation solutions for complex operational challenges. The ideal candidate performs deep data analysis to prioritize server repair automation in a hyperscale environment, drives solutions through code, and collaborates effectively with globally distributed teams through clear written communication.

Global Production Systems Engineer Responsibilities
  • Identify and root cause systemic issues across the server fleet and drive resolutions to maximize uptime and utilization by leveraging hardware failure data and diagnostic telemetry
  • Write, review, and maintain code for diagnostic and automation tooling that supports quality and efficient delivery of production servers at hyperscale
  • Own and develop diagnostic tooling requirements that enable frontline operations teams to efficiently manage and repair the server fleet
  • Drive the escalation process for Data Center Operations to identify, root cause, and resolve complex tooling and hardware issues affecting fleet health
  • Execute operational validation and verification activities for new product integration into the production environment
  • Collaborate with cross-functional tooling teams to provide an operations-centric perspective on open issues and contribute to their development roadmaps
  • Perform deep data analysis to prioritize automation opportunities for server repair workflows in a large-scale, heterogeneous hardware environment
  • Build cross-functional relationships and influence policies and procedures to improve global data center operations consistency and efficiency
  • Mentor other engineers on evaluating and resolving fleet issues and defining improvements to tools and operational processes
  • Travel up to 25% to support global data center operations and new site deployments
Minimum Qualifications
  • Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience
  • 6+ years of experience in production systems engineering, infrastructure engineering, or systems software development for large-scale hardware environments
  • 6+ years of experience with hardware lifecycle management, fleet automation, or data center operations systems spanning compute, storage, or networking infrastructure
  • Experience developing systems software or automation tooling in Python, Bash, PHP, C, or C++ for Linux-based production environments at scale
  • Experience with configuration and maintenance of production systems including web servers, load balancers, relational databases, storage systems, and messaging systems
  • Experience communicating technical designs and infrastructure decisions through written documentation and cross-functional stakeholder alignment across engineering and operations teams
Preferred Qualifications
  • Experience designing or operating configuration management and infrastructure-as-code systems for large heterogeneous hardware fleets
  • Experience supporting global, multi-site data center infrastructure deployments including hardware qualification and regional rollout coordination
  • Experience with data analysis and visualization tools used to prioritize fleet health initiatives and drive operational decision-making
  • Familiarity with distributed systems monitoring, alerting, and automated remediation pipelines at hyperscale
About Meta

Meta builds technologies that help people connect, find communities, and grow businesses. When Facebook launched in 2004, it changed the way people connect. Apps like Messenger, Instagram and WhatsApp further empowered billions around the world. Now, Meta is moving beyond 2D screens toward immersive experiences like augmented and virtual reality to help build the next evolution in social technology. People who choose to build their careers by building with us at Meta help shape a future that will take us beyond what digital connection makes possible today—beyond the constraints of screens, the limits of distance, and even the rules of physics.

Meta is proud to be an Equal Employment Opportunity and affirmative Action employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, or related medical conditions), sexual orientation, gender, gender identity, gender expression, transgender status, sexual stereotypes, age, status as a protected veteran, status as an individual with a disability, or other applicable legally protected characteristics. We also consider qualified applicants with criminal histories, consistent with applicable federal, state and local law. Meta participates in the E-Verify program in certain locations, as required by law. Please note that Meta may leverage artificial intelligence and machine learning technologies in connection with applications for employment.

Meta is committed to providing reasonable accommodations for candidates with disabilities in our recruiting process. If you need any assistance or accommodations due to a disability, please let us know at accommodations-ext@meta.com.

$144,000/year to $204,000/year + bonus + equity + benefits

Individual compensation is determined by skills, qualifications, experience, and location. Compensation details listed in this posting reflect the base hourly rate, monthly rate, or annual salary only, and do not include bonus, equity or sales incentives, if applicable. In addition to base compensation, Meta offers benefits. Learn more about benefits at Meta.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Center Production Operations Engineer
Data Center Production Operations Engineer

Meta • Bowling Green (OH)

On-site
USD 83,000 - 130,000
Equity
Benefits
Bonus
Data Center Production Operations Engineer
Data Center Production Operations Engineer

Meta • New Albany (OH)

On-site
USD 111,010 - 158,995
Data Center Production Operations Engineer
Data Center Production Operations Engineer

Meta • Hillsboro (OR)

On-site
USD 71,000 - 103,000
Data Center Production Operations Engineer
Data Center Production Operations Engineer

Meta • Huntsville (AL)

On-site
USD 84,000 - 130,000
Enterprise Systems Engineer
Enterprise Systems Engineer

Meta • Seattle (WA)

On-site
USD 142,000 - 206,000
Production Network Engineer
Production Network Engineer

Meta • United States

On-site
USD 162,000 - 227,000
Bonus
Equity
Benefits
Production Engineer, Network
Production Engineer, Network

Meta • Menlo Park (CA)

On-site
USD 154,000 - 217,000
Bonus
Equity
Benefits
Critical Operations Manager, Data Center Facilities
Critical Operations Manager, Data Center Facilities

Meta • Rayville (LA)

On-site
USD 168,000 - 234,000
Bonus
Equity
Benefits
Network Production Engineer, Infrastructure
Network Production Engineer, Infrastructure

Meta • Menlo Park (CA)

Hybrid
USD 122,000 - 181,000
Senior Data Center Facilities Engineer
Senior Data Center Facilities Engineer

Meta • Kuna (ID)

On-site