Production Systems Engineer, AI Systems

Meta

Austin (TX)

On-site

USD 144,000 - 204,000

Full time

18 hours ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Bonus
Equity
Benefits

Job summary

Meta is seeking a Hardware Systems Engineer to support the new product introduction of AI and high-performance computing infrastructure for large-scale data center deployments. You will collaborate with hardware design, firmware, software, networking, and capacity engineering teams to validate and scale AI hardware systems from bring-up to production readiness.

Responsibilities include leading end-to-end system validation for AI accelerators, driving bring-up and validation of AI server systems,

Qualifications

  • Bachelor's degree in CS, CE, or equivalent practical experience.
  • 6+ years in hardware systems engineering, silicon or firmware validation, or system bring-up for AI servers or accelerators.
  • Experience with ASIC bring-up, board-level debugging, or large-scale data center validation.
  • Experience writing test specs, validation procedures, and debug methodologies.
  • Experience leading root-cause analysis across hardware, firmware, and software.
  • Experience with PCIe, NVLink, DDR5, or HBM in AI/HPC validation.

Responsibilities

  • Lead end-to-end system validation strategies for AI and HPC hardware platforms in data centers.
  • Drive bring-up, characterization, and validation of AI server systems and components.
  • Develop and maintain test specifications, validation procedures, and debug guides.
  • Investigate root-cause failures across silicon, firmware, and software with cross-functional teams.
  • Triage hardware/firmware defects and track progress on NPI milestones.
  • Improve test coverage and tooling across the NPI lifecycle.
  • Define acceptance criteria and deployment readiness standards for new AI hardware.
  • Collect and analyze data to surface hardware quality trends for go/no-go decisions.
  • Communicate validation status and risks to internal teams and external vendors.
  • Collaborate with firmware/software to define telemetry and remote management interfaces.

Skills

Hardware systems engineering
Silicon validation
Firmware validation
System bring-up
Test specification development
Root-cause analysis
High-speed interconnects
Telemetry data analysis

Education

Bachelor's degree in CS/CE or equivalent

Tools

Python scripting
Linux servers

Job description

Meta is seeking a Hardware Systems Engineer to support the new product introduction (NPI) of next-generation AI and high-performance computing infrastructure for large-scale data center deployments. In this role, you will work at the intersection of server systems, AI applications and data center operations, partnering with hardware design, firmware, software, networking, and capacity engineering teams to validate and scale cutting-edge AI hardware systems from early bring-up through production readiness.

Production Systems Engineer, AI Systems Responsibilities:
  • Lead end-to-end system validation strategies for AI and HPC hardware platforms, including AI accelerators, GPU clusters, and high-bandwidth memory subsystems in data center environments
  • Drive hands-on bring-up, characterization, and validation of AI server systems and associated components such as PCIe, NVLink, DRAM, and high-speed networking fabrics
  • Develop and maintain test specifications, validation procedures, and debug guides tailored to AI infrastructure NPI programs
  • Investigate and root-cause complex system failures spanning silicon, firmware, software, and hardware layers in collaboration with cross-functional engineering teams
  • Triage and track hardware and firmware defects through resolution while maintaining forward progress on NPI program milestones
  • Identify gaps in test coverage and drive improvements to test methodologies, tooling, and automation frameworks across the NPI lifecycle
  • Partner with AI platform and capacity engineering teams to define acceptance criteria and deployment readiness standards for new AI hardware systems
  • Guide data collection, analysis, and reporting efforts to surface systemic hardware quality trends and inform go/no-go decisions for production deployment
  • Communicate validation status, risk assessments, and technical findings to internal engineering teams and external hardware vendors
  • Collaborate with firmware and software teams to define hardware-software interface requirements for telemetry, diagnostics, and remote management of AI infrastructure
Minimum Qualifications:
  • Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience
  • 6+ years of experience in hardware systems engineering, silicon validation, firmware validation, or system-level bring-up for AI servers, GPUs, TPUs, or AI accelerator platforms
  • Experience in one or more of the following domains: ASIC bring-up and characterization, board-level debug, firmware validation, or large-scale system validation in data center environments
  • Experience developing test specifications, validation procedures, and debug methodologies for complex hardware systems
  • Experience leading root-cause analysis and troubleshooting of system-level failures across hardware, firmware, and software stacks
  • Experience with high-speed interconnects or memory subsystems such as PCIe, NVLink, DDR5, or HBM in the context of AI or HPC system validation
  • Experience analyzing system telemetry and fleet health data to identify reliability trends and drive engineering improvements
Preferred Qualifications:
  • Proficiency in scripting or programming languages such as Python for automation of infrastructure workflows and data analysis
  • Familiarity with Linux-based server environments and data center management tooling used in large-scale production operations
  • Experience defining hardware-software interface requirements for telemetry, out-of-band management, or remote diagnostics in data center AI systems
  • Experience with high-speed interconnects and memory subsystems such as PCIe, NVLink, InfiniBand, DDR5, or HBM in the context of AI or HPC infrastructure operations
About Meta:

Meta builds technologies that help people connect, find communities, and grow businesses. When Facebook launched in 2004, it changed the way people connect. Apps like Messenger, Instagram and WhatsApp further empowered billions around the world. Now, Meta is moving beyond 2D screens toward immersive experiences like augmented and virtual reality to help build the next evolution in social technology. People who choose to build their careers by building with us at Meta help shape a future that will take us beyond what digital connection makes possible today—beyond the constraints of screens, the limits of distance, and even the rules of physics.

Meta is proud to be an Equal Employment Opportunity and Aff… We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, or related medical conditions), sexual orientation, gender, gender identity, gender expression, transgender status, sexual stereotypes, age, status as a protected veteran, status as an individual with a disability, or other applicable legally protected characteristics. We also consider qualified applicants with criminal histories, consistent with applicable federal, state and local law. Meta participates in the E-Verify program in certain locations, as required by law. Please note that Meta may leverage artificial intelligence and machine learning technologies in connection with applications for employment.

Meta is committed to providing reasonable accommodations for candidates with disabilities in our recruiting process. If you need any assistance or accommodations due to a disability, please let us know at accommodations-ext@meta.com.

$144,000/year to $204,000/year + bonus + equity + benefits

Individual compensation is determined by skills, qualifications, experience, and location. Compensation details listed in this posting reflect the base hourly rate, monthly rate, or annual salary only, and do not include bonus, equity or sales incentives, if applicable. In addition to base compensation, Meta offers benefits. Learn more about benefits at Meta.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI/HPC System Performance Engineer
AI/HPC System Performance Engineer

Meta • Menlo Park (CA)

On-site
USD 154,000 - 217,000
Bonus
Equity
Benefits
Production Systems Engineer, AI Systems
Production Systems Engineer, AI Systems

Meta • Menlo Park (CA)

On-site
USD 173,000 - 245,000
Bonus
Equity
Benefits
Product Quality Engineer
Product Quality Engineer

Meta • Bellevue (WA)

On-site
USD 144,000 - 204,000
Bonus
Equity
Benefits
Global Production Systems Engineer
Global Production Systems Engineer

Meta • Nebraska

On-site
USD 144,000 - 204,000
Product Quality Engineer
Product Quality Engineer

Meta • Austin (TX)

On-site
USD 144,000 - 204,000
Software Engineer, Systems ML Engineering
Software Engineer, Systems ML Engineering

Meta • Menlo Park (CA)

On-site
USD 183,000 - 257,000
Software Engineering Manager - Neural Interface ML Infra
Software Engineering Manager - Neural Interface ML Infra

Meta • New York (NY)

On-site
USD 184,000 - 257,000
Bonus
Equity
Benefits
Data Center Production Operations Engineer
Data Center Production Operations Engineer

Meta • Bowling Green (OH)

On-site
USD 83,000 - 130,000
Equity
Benefits
Bonus
Software Engineering Manager - Neural Interface ML Infra
Software Engineering Manager - Neural Interface ML Infra

Meta • Redmond (WA)

On-site
USD 184,000 - 257,000
Bonus
Equity
Benefits
Research Scientist, AI & Systems Co-Design (PhD)
Research Scientist, AI & Systems Co-Design (PhD)

Meta • Menlo Park (CA)

On-site
USD 122,000 - 181,000