AI/HPC Performance Engineer: Scale Large AI Clusters

Meta

Menlo Park (CA)

On-site

USD 154,000 - 217,000

Full time

9 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Bonus
Equity
Benefits

Job summary

Meta is seeking an AI/HPC System Performance Engineer to drive end-to-end performance characterization of large-scale AI training and inference clusters. You will profile workloads, build dashboards, and resolve performance regressions across RDMA fabrics, NCCL/MPI, and network infrastructure.

The role sits on the Network Infrastructure Engineering team and collaborates with hardware and AI research groups.

Qualifications

  • 5+ years coding experience in C, C++, Python or similar languages.
  • Experience profiling and optimizing distributed AI or HPC workloads with GPU interconnects, RDMA, NCCL/MPI.
  • Experience debugging complex performance issues across multi-layer systems.

Responsibilities

  • Profile and benchmark AI training/inference workloads on large-scale HPC clusters to identify bottlenecks.
  • Develop performance analysis frameworks and dashboards for system-level metrics (GPU utilization, bandwidth, latency).
  • Investigate performance regressions in distributed AI/HPC environments including RDMA fabrics and scheduling.
  • Collaborate with network, hardware, and AI teams to validate new HPC configurations.
  • Design and execute capacity/scalability experiments to inform network topology decisions.

Skills

C/C++/Python
Performance profiling
Distributed AI workloads
Cross-functional work

Education

Bachelor's degree in CS/CE/related field

Tools

NCCL
MPI
RDMA networking
PyTorch

Job description

Meta is seeking an AI/HPC System Performance Engineer to drive end-to-end performance characterization of large-scale AI training and inference clusters. You will profile workloads, build dashboards, and resolve performance regressions across RDMA fabrics, NCCL/MPI, and network infrastructure.

The role sits on the Network Infrastructure Engineering team and collaborates with hardware and AI research groups.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Infrastructure Engineer
Senior AI Infrastructure Engineer

AI Breaking Wire • Menlo Park (CA), Northern (KY)

Hybrid
USD 200,000 - 350,000
RSUs
Health benefits
Parental leave
+1
Senior HPC-AI Cluster Architect (Equity)
Senior HPC-AI Cluster Architect (Equity)

NVIDIA • Santa Clara (CA)

On-site
USD 176,000 - 334,000
Equity
Benefits
AI/HPC Cluster Architect
AI/HPC Cluster Architect

AMD • United States

On-site
USD 140,000 - 210,000
AI/HPC Cluster Architect
AI/HPC Cluster Architect

AMD • Austin (TX)

On-site
USD 140,000 - 200,000
AI/HPC System Performance Engineer
AI/HPC System Performance Engineer

Meta • Menlo Park (CA)

On-site
USD 154,000 - 217,000
Bonus
Equity
Benefits
Distributed ML Systems Engineer - Scaling (Equity)
Distributed ML Systems Engineer - Scaling (Equity)

Meta • Menlo Park (CA)

On-site
USD 154,000 - 217,000
Bonus
Equity
Benefits
AI/HPC Cluster Architect
AI/HPC Cluster Architect

Advanced Micro Devices • Austin (TX)

On-site
USD 120,000 - 180,000
AMD benefits
Senior HPC-AI Systems Architect (Equity)
Senior HPC-AI Systems Architect (Equity)

Nvidia Corporation in • Santa Clara (CA)

On-site
USD 176,000 - 334,000
AI/HPC Cluster Architect — Data Center Power & Network
AI/HPC Cluster Architect — Data Center Power & Network

AMD • Austin (TX)

On-site
USD 120,000 - 180,000
AMD benefits at a glance
Senior AI Infrastructure Engineer — Scale GPU Clusters Remote
Senior AI Infrastructure Engineer — Scale GPU Clusters Remote

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 280,000 - 420,000
Equity options
Health, vision, dental benefits
Unlimited PTO
+2