Production Systems Engineer, Sustaining

Meta

Austin (TX)

On-site

USD 160,000 - 230,000

Full time

48 hours ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Meta seeks an experienced Production Systems Engineer to join the Release to Production (RTP) team. You will work on hardware lifecycle, testing, and production readiness for servers across data centers, collaborating with hardware designers, vendors, and data center operations.

The role emphasizes developing scalable testing practices, diagnosing issues, and driving improvements to test quality and workflows for AI/HPC workloads at scale.

Qualifications

  • Bachelor's degree in a technical field or equivalent practical experience.
  • 6+ years hardware systems tech or scale production support.
  • Experience deploying and productionizing CPU systems and related components.
  • Experience in software and hardware co-design for hyperscale systems.

Responsibilities

  • Develop robust, scalable, reliable practices for AI and HPC infrastructure at scale.
  • Interface with vendors and internal teams to understand system architecture and testing.
  • Create experiments and tooling to diagnose hardware/firmware/software health issues.
  • Implement sustaining workflows across hardware and software stacks and communicate them internally.
  • Troubleshoot and root-cause system failures with stakeholders.
  • Drive discussions on test specs and methodologies to improve test quality.

Skills

Hardware knowledge
Production support
Programming (Python/C/C++)
System debugging
Cross-functional collaboration

Education

Bachelor's degree in CS/Engineering

Tools

NCCL
PyTorch
CUDA
VM environments

Job description

About

Meta is seeking an experienced Production Systems Engineer to join our Release to Production (RTP) team. Our servers and data centers are the foundation upon which our rapidly scaling infrastructure operates efficiently to deliver our innovative services. The RTP team is responsible for the Hardware Lifecycle of all Meta servers including pre-production hands‑on system and hardware debugging and stress testing, enabling production‑ready system monitoring, automated provisioning and automated remediation of issues. RTP Engineers work closely with hardware designers, system manufacturers, component vendors, capacity engineering, production engineering, Meta services, and data center operations teams to test systems before release to our production data centers, and to track the health and life cycle of servers in production.

Responsibilities
  • Develop robust, scalable, reliable practices for supporting AI and HPC infrastructure at scale
  • Interface with external vendors and internal hardware, mechanical, power, thermal, manufacturing and software engineers to understand system architecture to develop and execute the test suites for various architectures
  • Proactively create experiments and tooling to detect and diagnose hardware/firmware/software health issues
  • Implement sustaining workflows across hardware and software stacks, develop and communicate sustaining practices internally
  • Troubleshoot, diagnose and root cause of system failures and isolate components and failure scenarios while working with internal and external stakeholders
  • Drive necessary discussions with external and internal teams on test specification and methodologies to improve test quality on an ongoing basis
Minimum Qualifications
  • Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience
  • 6+ years of experience in hardware systems technologies or supporting production hardware at scale
  • Experience in deploying and productionizing CPU systems and/or related components at scale
  • Experience in software and hardware co-design for hyperscale systems
  • Experience in object oriented programming (e.g., Python, C/C++)
  • Experience in different server and network datacenter systems 8+ years of experience in working with CPU systems, including hardware and software components, co-design
  • 8+ years of experience in providing production support for large-scale systems
  • Knowledge of AI/HPC hardware requirements and specifications (e.g., configuring hardware components, GPU, memory, network for AI/HPC workloads)
  • Experience in developing or debugging CPU systems in a VM environment, performance optimizations, including familiarity with relevant tools, libraries, and frameworks (e.g., NCCL, PyTorch, CUDA)
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Production Systems Engineer, Sustaining
Production Systems Engineer, Sustaining

Meta Careers • Austin (TX), Menlo Park (CA)

On-site
USD 170,000 - 250,000
Production Systems Engineer
Production Systems Engineer

Meta • Menlo Park (CA)

On-site
USD 180,000 - 240,000
Production Systems Engineer, NPI Storage
Production Systems Engineer, NPI Storage

Meta • Austin (TX)

On-site
USD 180,000 - 260,000
Senior Production Systems Engineer AI/HPC Infra
Senior Production Systems Engineer AI/HPC Infra

Meta • Austin (TX)

On-site
USD 160,000 - 230,000
Cloud Production Platform Engineer
Cloud Production Platform Engineer

Meta • Los Lunas (NM)

On-site
USD 150,000 - 230,000
Global Production Platform Engineer
Global Production Platform Engineer

Meta • Menlo Park (CA)

On-site
USD 140,000 - 210,000
Production Systems Engineer, Tooling
Production Systems Engineer, Tooling

Meta • Menlo Park (CA)

On-site
USD 120,000 - 150,000
Production Systems Engineer, AI Systems
Production Systems Engineer, AI Systems

Meta • Austin (TX)

On-site
USD 144,000 - 204,000
Production Systems Engineer, AI Network
Production Systems Engineer, AI Network

Meta • Menlo Park (CA)

On-site
USD 144,000 - 204,000
Production Systems Engineer, AI Systems
Production Systems Engineer, AI Systems

Meta • Menlo Park (CA)

On-site
USD 173,000 - 245,000