Software Engineering Manager, AI Systems & HPC

Takadam

Menlo Park (CA)

On-site

USD 177,000 - 251,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health insurance
Paid leave
Professional development
Equity

Job summary

Meta is seeking an Software Engineering Manager for the AI Systems Co-Design team, based in Menlo Park. You will lead HPC and AI infrastructure efforts at the intersection of hardware and software, delivering compute systems for large-scale models like LLaMA and DLRM.

You’ll collaborate with academia and industry, drive performance optimization, and guide cross-functional teams across engineering, research, and product groups to push the boundaries of scalable AI infrastructure.

Qualifications

  • Proven leadership in high-performance computing (HPC) and AI/ML systems.
  • Experience with communication libraries (NCCL, RCCL, UCC, MPI).
  • Background in GPU/ASIC kernel development (CUDA, ROCm).
  • Skilled in distributed systems, systems architecture, and performance optimization.
  • Experience managing large-scale engineering programs and cross-functional collaboration.

Responsibilities

  • Lead the AI Systems Co-Design team, aligning hardware, software, and AI infra.
  • Deliver high-performance compute systems to support large-scale AI models like LLaMA and DLRM.
  • Collaborate with top academic institutions and industry partners to advance infrastructure.
  • Shape Meta’s AI infrastructure strategy through cross-functional leadership.

Skills

HPC leadership
AI/ML systems
Distributed systems
Systems architecture
Performance optimization
Cross-functional collaboration
Large-scale program management
Team leadership

Tools

CUDA
ROCm
NCCL
RCCL
UCC
MPI

Job description

Meta is seeking an Software Engineering Manager for the AI Systems Co-Design team, based in Menlo Park. You will lead HPC and AI infrastructure efforts at the intersection of hardware and software, delivering compute systems for large-scale models like LLaMA and DLRM.

You’ll collaborate with academia and industry, drive performance optimization, and guide cross-functional teams across engineering, research, and product groups to push the boundaries of scalable AI infrastructure.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineering Manager – AI Systems Co-Design at Meta
Software Engineering Manager – AI Systems Co-Design at Meta

Takadam • Menlo Park (CA)

On-site
USD 177,000 - 251,000
Health insurance
Paid leave
Professional development
+1
Software Engineering Manager - AI Strategy & Scale
Software Engineering Manager - AI Strategy & Scale

Meta • Bellevue (WA)

On-site
USD 180,000 - 260,000
Staff ML Systems Engineer - Scalable AI Infrastructure
Staff ML Systems Engineer - Scalable AI Infrastructure

Meta • Menlo Park (CA)

On-site
USD 183,000 - 257,000
Engineering Manager – AI Infra & ML Systems
Engineering Manager – AI Infra & ML Systems

Meta • New York (NY)

On-site
USD 184,000 - 257,000
Bonus
Equity
Benefits
Engineering Manager, Neural AI Infrastructure & Hardware
Engineering Manager, Neural AI Infrastructure & Hardware

Meta • Bellevue (WA)

On-site
USD 184,000 - 257,000
Lead AI Systems Engineer - Production ML & Platforms
Lead AI Systems Engineer - Production ML & Platforms

Meta • Menlo Park (CA)

On-site
USD 219,000 - 301,000
Neural AI Infra Engineering Manager
Neural AI Infra Engineering Manager

Meta • Burlingame (CA)

On-site
USD 184,000 - 257,000
Bonus
Equity
Benefits
Engineering Manager, Neural Interface ML Infra
Engineering Manager, Neural Interface ML Infra

Meta • Redmond (WA)

On-site
USD 184,000 - 257,000
Bonus
Equity
Benefits
ML Systems Engineer: AI Infra & GPU Acceleration
ML Systems Engineer: AI Infra & GPU Acceleration

Meta • San Francisco (CA)

On-site
USD 180,000 - 240,000
Bonus
Equity
Software Engineering Manager
Software Engineering Manager

Meta • Bellevue (WA)

On-site
USD 180,000 - 260,000