Senior Production Systems Engineer AI Infrastructure

Meta

Austin (TX)

On-site

USD 140,000 - 200,000

Full time

9 days ago
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Meta is seeking a Systems Engineer to join the RTP-NPI team, focusing on AI/ML initiatives and large-scale training/inference. You will work on end-to-end hardware lifecycle, prototyping, debugging, and production readiness for datacenter-scale AI systems in Austin.

The role involves integrating scale-up and scale-out interfaces, diagnosing issues, and collaborating with HW/SW co-designers and data center teams to improve platform quality and reliability.

Qualifications

  • Bachelor's degree in CS/Engineering or equivalent practical experience.
  • 6+ years in AI/ML/HPC network evaluation, NICs, or RDMA.
  • Strong TCP/IP knowledge; experience with iperf.
  • Experience with Linux and server hardware components.
  • Hands-on troubleshooting and debugging; Python scripting.
  • Experience with large-scale deployments and RoCE networks.

Responsibilities

  • Lead integration of scale-up interfaces (NVLink, XGMI, RoCE) and scale-out NICs for AI platforms.
  • Understand AI workloads and communicate patterns for NPI.
  • Create experiments and tooling to detect and diagnose hardware/firmware/software issues.
  • Contribute to hacks and explorations of future AI technologies.
  • Troubleshoot and root-cause system failures with internal and external partners.
  • Develop visibility through data visualization to address hardware health issues.
  • Leverage production experience to drive improvements in product quality.

Skills

Networking basics
TCP/IP
Python scripting
Linux
Debugging
RDMA/RoCE
NICs
AI servers

Education

Bachelor's degree in CS/Engineering
Bachelor’s degree in Engineering or Computer Science

Tools

iperf

Job description

Meta is seeking a Systems Engineer to join the RTP-NPI team, focusing on AI/ML initiatives and large-scale training/inference. You will work on end-to-end hardware lifecycle, prototyping, debugging, and production readiness for datacenter-scale AI systems in Austin.

The role involves integrating scale-up and scale-out interfaces, diagnosing issues, and collaborating with HW/SW co-designers and data center teams to improve platform quality and reliability.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Systems Production Engineer – HPC & Data Center Validation
AI Systems Production Engineer – HPC & Data Center Validation

Meta • Austin (TX)

On-site
USD 144,000 - 204,000
Production Systems Engineer, Fleet AI Systems
Production Systems Engineer, Fleet AI Systems

Meta • Austin (TX)

On-site
USD 140,000 - 200,000
NPI AI Hardware Validation Engineer
NPI AI Hardware Validation Engineer

Meta • Menlo Park (CA)

On-site
USD 118,000 - 170,000
Lead AI Systems Engineer - Production ML & Platforms
Lead AI Systems Engineer - Production ML & Platforms

Meta • Menlo Park (CA)

On-site
USD 219,000 - 301,000
Senior AI Infrastructure Engineer
Senior AI Infrastructure Engineer

AI Breaking Wire • Menlo Park (CA), Northern (KY)

Hybrid
USD 200,000 - 350,000
RSUs
Health benefits
Parental leave
+1
AI Infrastructure TPM: Scale Hardware for AI Systems
AI Infrastructure TPM: Scale Hardware for AI Systems

Meta • Menlo Park (CA)

On-site
USD 168,000 - 234,000
Senior Production Engineer: Scale & AI-Driven Backend
Senior Production Engineer: Scale & AI-Driven Backend

Meta • Nashville (TN)

On-site
USD 184,000 - 257,000
Senior Systems ML Engineer — High-Performance AI Infra
Senior Systems ML Engineer — High-Performance AI Infra

Meta • Annapolis (MD)

On-site
USD 154,000 - 217,000
Equity
Benefits
Senior Systems ML Engineer - Scalable AI Infra
Senior Systems ML Engineer - Scalable AI Infra

Meta • Raleigh (NC)

On-site
USD 154,000 - 217,000
Production Systems Engineer, AI Systems
Production Systems Engineer, AI Systems

Meta • Austin (TX)

On-site
USD 144,000 - 204,000