A complete application in a minute — tailored resume and cover letter, ready to send.
Meta is seeking an experienced hardware engineering professional to lead data center compute, storage, and GPU platforms across Meta-owned data centers and cloud capacity. You will diagnose systemic hardware issues, coordinate with hardware design, tooling, and onsite teams, and ensure platform reliability for production workloads.
The role requires a strong background in server hardware, Linux, and cross-functional collaboration, with a focus on scalable infrastructure and proactive problem
Meta is seeking a experienced hardware engineering professional with technical leadership experience and skills in Server Hardware, high performance Compute, Storage, Accelerators (GPU), and/or Networking, ideally in a data center environment. Our data centers, the global fleet of servers installed in them, and the growing cloud capacity Meta operates alongside them are the foundation upon which our rapidly scaling infrastructure efficiently operates and on which our innovative services are delivered. Our Production Platform Engineers are responsible for the performance of the compute, storage, and accelerator (GPU) platforms that run Meta's production workloads, both in our own data centers and on capacity Meta consumes from cloud providers. They are responsible for the health of the server fleet, from NPI production operational test through end of life. Responsibilities include identifying systemic hardware, firmware, and tooling issues; engaging in hands on problem solving; and collaborating effectively with engineering, tooling, and onsite provisioning and break fix teams to improve performance of the fleet.This position carries the cloud platform scope for the team. Meta's cloud footprint is growing toward the scale of the fleet we run ourselves, across several cloud providers, and product groups expect the same reliability and availability from it. Because Meta does not design this hardware and does not have physical access to it, the engineer in this role works through each cloud provider's own repair and escalation process, reaches diagnostic conclusions with partial telemetry, and builds the technical evidence that supports cloud provider accountability. This remains a platform engineering role across both environments. It is not a cloud only role.This position requires a high degree of understanding of our platforms, leveraging technical expertise and using data to identify trends across the fleet and proactively identify and mitigate performance and reliability issues. Experience organizing and prioritizing work across multiple teams in a large-scale, distributed environment and experience partnering with cross-functional teams to drive outcomes in a large-scale, distributed environment.
Cloud Production Platform Engineer Responsibilities:
$173,000/year to $245,000/year + bonus + equity + benefits
Internet
Meta is proud to be an Equal Employment Opportunity and Affiantive Action employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, or related medical conditions), sexual orientation, gender, gender identity, gender expression, transgender status, sexual stereotypes, age, status as a protected veteran, status as an individual with a disability or other applicable legally protected characteristics. We also consider qualified applicants with criminal histories, consistent with applicable federal, state and local law. Meta participates in the E-Verify program in certain locations, as required by law. Please note that Meta may leverage artificial intelligence and machine learning technologies in connection with applications for employment.
Meta is committed to providing reasonable accommodations for candidates with disabilities in our recruiting process. If you need any assistance or accommodations due to a disability, please let us know at accommodations-ext@meta.com.