Data Center Production Operations Engineer

Meta

El Paso (TX)

On-site

USD 84,000 - 130,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Meta is seeking a Production Operations Engineer to join our Data Centers. You will maintain server hardware and Linux systems in a fast-paced, scalable environment, collaborating with cross-functional teams to improve reliability.

This role includes on-call rotation, root-cause analysis, automation, and mentoring teammates, with travel up to 15% and a focus on data-driven improvements to keep services running at peak uptime.

Qualifications

  • BS, BA or BEng in technical field or commensurate experience.
  • 5+ years of technical IT experience in infrastructure (SysAdmin/DevOps/SRE).
  • Linux in a complex IT environment; triage, debug, and troubleshoot server issues.
  • Hands-on server hardware knowledge incl. storage.
  • Interdependencies of data center functions: electrical, cooling, cabling, security, network.
  • Experience driving to root cause of technical issues.
  • Experience in tech projects related to process, technology or automation.
  • Effective communicator; tailor messages to audience.
  • Knowledge of HTTP, DNS, RAID, DHCP.
  • Experience guiding external vendors.
  • Scripting languages Bash/Python/SQL/Rust/Go/Perl.
  • Knowledge of IPMI / serial console.
  • Experience using data/metrics to drive decisions.

Responsibilities

  • Resolve tickets and address root cause through remote and physical checks.
  • Lead RCAs for hardware, automation, and network issues.
  • Collaborate with cross-functional teams on hardware, tooling, and processes.
  • Lead introduction of new platforms/hardware with partners.
  • Use data to identify issues and communicate with stakeholders.
  • Identify corrective actions with internal teams and vendors.
  • Influence future design changes for serviceability.
  • Solve hardware/software issues at scale via scripting and tooling.
  • Pursue continuous process/tool improvements.
  • Use data analytics to maximize server uptime and utilization.
  • Mentor and train teammates on resolution approaches.
  • Provide engineering support to leadership and teams.
  • Maintain runbooks and documentation.
  • Build cross-functional relationships to improve ops.
  • Participate in 24/7 on-call rotation.
  • Travel up to 15%.

Skills

Linux
Server hardware
Scripting
Root cause analysis
Data center ops
Communication
Automation
Networking
IPMI
Vendor management
Troubleshooting
Data analytics

Education

Bachelor's degree in technical field

Tools

IPMI
Serial console

Job description

Summary

Meta is seeking a forward thinking experienced engineer to join the Production Operations team within our Data Centers. These Data Centers are the foundation upon which our rapidly scaling infrastructure efficiently operates and upon which our innovative services are delivered. Meta is at the leading edge of the global data center industry both in terms of how data centers are designed and operated. This person should enjoy working in a fast paced, technical environment where adaptability and flexibility will be key to their success. We seek an IT professional with advanced, hands-on technical skills in server hardware and Linux - ideally in a Data Center environment. Having broad knowledge of server administration and participating in projects in a large-scale distributed data center environment is a core competency of this individual. The candidate should also have working knowledge and experience in a few of the following core areas: Hardware repair, OS management, Tooling and Automation, Networking, or Technical Project Management.

Required Skills

Data Center Production Operations Engineer Responsibilities:

  1. Support platform health by successfully resolving and closing tickets, while addressing the overall issue (i.e. addressing root cause) including, but not limited to, remote troubleshooting and physical inspection of services in data halls

  2. Participate in deep dives and root cause analysis of highly technical issues within the data center, ranging from automated tooling to hardware failures and network issues

  3. Collaborate with cross-functional teams on projects and initiatives related to topics such as process, hardware and automation

  4. Point of contact for the introduction of new platforms and hardware to the site, in collaboration with partners and global resources, accelerating the time it takes to bring these products to sustained mass production

  5. Use tools and data analysis effectively to identify issues. Take actions to communicate with all stakeholders appropriately and manage or escal… as needed

  6. Identify corrective actions of hardware issues, work with internal teams and vendors

  7. influence future design changes to ensure ease of serviceability

  8. Solve systemic hardware and/or software issues at scale using scripting, automation, and tooling to drive global resolution

  9. Continuously evaluate and identify areas for improvement in processes, tools, and systems to optimize efficiency and quality of repairs

  10. Use data analytics to drive maximum server up-time and utilization rates, understanding hardware failure rates and service level agreements

  11. Support and train team members to evaluate and identify better ways to resolve issues, and define updates to tools and processes

  12. Provide engineering support and be a go-to technical resource for the team, leadership, and cross-functional teams in operating and maintaining data center servers

  13. Maintain and update documentation i.e. procedures, runbooks and guides

  14. Build cross functional relationships and influence policies and procedures that improve global data center operations

  15. Participate in 24/7 on-call rotation

  16. Ability to travel up to 15% of the time

Minimum Qualifications
  1. BS, BA or BEng in technical field or commensurate experience

  2. 5+ years of technical IT experience within an infrastructure environment, in a role such as Systems Administrator, DevOps Engineer, or Site Reliability Engineer

  3. Intermediate-level understanding in Linux (or equivalent OS) in a complex IT environment with the ability to triage, debug, and troubleshoot server issues

  4. Hands-on experience and knowledge of server hardware and components, including storage

  5. Intermediate-level knowledge of the interdependencies of data center functions and technologies including electrical, cooling, structured cabling, security, and network

  6. Experience managing technical issues and driving to the root cause

  7. Experience participating in technical projects related to areas such as process improvement, technology, and/or automation

  8. Ability to communicate effectively, in a clear and concise manner, appropriately tailoring messages to the audience

  9. Intermediate-level knowledge of technologies such as HTTP, DNS, RAID, and DHCP

  10. Experience in providing technical guidance to external vendors

  11. Experience in debugging, modifying and developing commonly used scripting or programming languages in at least one of these languages: Bash, PHP, Python, SQL, Rust, Go or Perl

  12. Knowledge of out-of-band/lights-out server communication methods, such as IPMI and serial console

  13. Experience using data and metrics to drive decisions

Preferred Qualifications
  1. Experience in fostering growth in others, and driving influence across all organizational levels

  2. Experience in a large-scale data center environment

  3. Experience with large-scale AI implementations

  4. Six Sigma knowledge/certification

Public Compensation

$83,990/year to $130,000/year + bonus + equity + benefits

Industry

Internet

Equal Opportunity

Meta is proud to be an Equal Employment Opportunity and Affirmative Action employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, or related medical conditions), sexual orientation, gender, gender identity, gender expression, transgender status, sexual stereotypes, age, status as a protected veteran, status as an individual with a disability, or other applicable legally protected characteristics. We also consider qualified applicants with criminal histories, consistent with applicable federal, state and local law. Meta participates in the E-Verify program in certain locations, as required by law. Please note that Meta may leverage artificial intelligence and machine learning technologies in connection with applications for employment.

Meta is committed to providing reasonable accommodations for candidates with disabilities in our recruiting process. If you need any assistance or accommodations due to a disability, please let us know at accommodations-ext@meta.com.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Center Production Operations Engineer
Data Center Production Operations Engineer

Meta • Bowling Green (OH)

On-site
USD 83,000 - 130,000
Equity
Benefits
Bonus
Data Center Production Operations Engineer
Data Center Production Operations Engineer

Meta • Huntsville (AL)

On-site
USD 84,000 - 130,000
Data Center Production Operations Engineer
Data Center Production Operations Engineer

Meta • Eagle Mountain (UT)

On-site
USD 71,000 - 103,000
Global Production Systems Engineer
Global Production Systems Engineer

Meta • Nebraska

On-site
USD 144,000 - 204,000
Data Center Production Operations Engineer
Data Center Production Operations Engineer

Meta • Oregon (WI)

On-site
USD 71,000 - 103,000
Bonus
Equity
Benefits
Lead Building Engineer
Lead Building Engineer

Meta • Montgomery (AL)

On-site
USD 129,002 - 186,000
Bonus
Equity
Health benefits
Network Engineer, Operations & Support
Network Engineer, Operations & Support

Meta • Santa Clara (CA)

On-site
USD 106,000 - 159,000
Bonus
Equity
Benefits
Head of Critical Data Center Operations
Head of Critical Data Center Operations

Meta • Temple (TX)

On-site
Critical Facility Engineer
Critical Facility Engineer

Meta • Eagle Mountain (UT)

On-site
USD 105,000 - 151,000
Data Center Lease Development Manager
Data Center Lease Development Manager

Meta • Carson City (NV)

On-site
USD 202,000 - 273,000
Bonus
Equity
Benefits