Senior AI Networking Engineer - NCCL & Multi-GPU Training

Meta Careers

Menlo Park, Northern (CA, KY)

Hybrid

USD 150,000 - 210,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Meta Careers in Menlo Park, CA is seeking a Software Engineer for SystemML - AI Networking to help build and optimize the NCCL-based communication stack used in large-scale GPU training. The role emphasizes GenAI/LLM scaling, reliability, and performance across the distributed AI network.

You will lead critical software projects, contribute to the PyTorch ecosystem, and collaborate with a team driving meta-wide ML innovations on a multi-GPU, multi-node platform.

Qualifications

  • Bachelor's degree or equivalent practical experience in Computer Science, Computer Engineering, or a related field.
  • Proven C/C++ and Python programming skills.
  • Proven track record of leading successful projects.

Responsibilities

  • Lead the development of the collective communication library for Meta's large-scale GPU training infrastructure, focusing on GenAI/LLM scaling.
  • Collaborate with cross-functional teams to optimize performance and reliability of distributed AI training.
  • Drive design decisions, code quality, and timely delivery of critical software components.

Skills

C/C++
Python
Project leadership

Education

Bachelor's degree in Computer Science or related field

Job description

Meta Careers in Menlo Park, CA is seeking a Software Engineer for SystemML - AI Networking to help build and optimize the NCCL-based communication stack used in large-scale GPU training. The role emphasizes GenAI/LLM scaling, reliability, and performance across the distributed AI network.

You will lead critical software projects, contribute to the PyTorch ecosystem, and collaborate with a team driving meta-wide ML innovations on a multi-GPU, multi-node platform.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Software Engineer - AI Networking & Distributed GPU Systems
Software Engineer - AI Networking & Distributed GPU Systems

Meta • Menlo Park (CA)

On-site
USD 154,000 - 217,000
Bonus
Equity
Benefits
Software Engineer, SystemML - AI Networking
Software Engineer, SystemML - AI Networking

Meta • Menlo Park (CA)

On-site
USD 154,000 - 217,000
Bonus
Equity
Benefits
AI/ML Network Infrastructure Engineer I
AI/ML Network Infrastructure Engineer I

Amazon • Cupertino (CA)

On-site
USD 127,000 - 185,000
Remote Senior AI Frameworks Communications Engineer
Remote Senior AI Frameworks Communications Engineer

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 160,000 - 280,000
Equity
Benefits
Senior AI Factory Deployment Architect (Multi-GPU)
Senior AI Factory Deployment Architect (Multi-GPU)

NVIDIA • Virginia (MN)

On-site
USD 148,000 - 235,750
Senior AI GPU Infrastructure Architect
Senior AI GPU Infrastructure Architect

ECLARO • Costa Mesa (CA)

On-site
USD 166,000 - 220,000
Senior AI Networking & Performance Engineer
Senior AI Networking & Performance Engineer

NVIDIA • Town of Texas (WI)

On-site
USD 272,000 - 431,250
Senior AI Infrastructure Lead: GPU Clusters & LLMs
Senior AI Infrastructure Lead: GPU Clusters & LLMs

Cadence Design Systems • San Jose (CA)

On-site
USD 137,000 - 254,000
Senior AI Networking and Performance Engineer for LLMs
Senior AI Networking and Performance Engineer for LLMs

Nvidia Corporation • Santa Clara (CA)

On-site
USD 272,000 - 431,250
Equity
Comprehensive benefits
Realtime ML Systems Engineer, Networking & AIOps
Realtime ML Systems Engineer, Networking & AIOps

NVIDIA Corporation • Santa Clara (CA), Northern (KY)

Hybrid
USD 152,000 - 288,000