Software Engineer - AI Networking & Distributed GPU Systems

Meta

Menlo Park (CA)

On-site

USD 154,000 - 217,000

Full time

10 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Bonus
Equity
Benefits

Job summary

Meta is seeking a Software Engineer for the AI Networking Software team, focused on NCCL-based distributed GPU communication for large-scale ML workloads. You will help design, optimize, and own software across multi-GPU and multi-node training stacks, enabling GenAI/LLM scaling with high performance and reliability.

You will lead technical initiatives, collaborate with ML engineers and infrastructure teams, and contribute to performance benchmarks and optimizations across CUDA and PyTorch

Qualifications

  • Bachelor's degree in Computer Science, Computer Engineering, or equivalent practical experience.
  • Proven C/C++ and Python programming skills.
  • Proven track record of leading successful projects.

Responsibilities

  • Provide technical leadership for the collective communication library development on Meta's large-scale GPU training infra with a focus on GenAI/LLM scaling.
  • Lead cross-functional technical projects and communicate decisions to technical and non-technical stakeholders.
  • Collaborate to optimize performance and reliability of the AI networking software stack across NCCL and PyTorch integration.

Skills

C/C++
Python
Leadership
Distributed systems

Education

Bachelor's degree in CS/CE
PhD in CS/CE (preferred)

Tools

NCCL
CUDA
PyTorch

Job description

Meta is seeking a Software Engineer for the AI Networking Software team, focused on NCCL-based distributed GPU communication for large-scale ML workloads. You will help design, optimize, and own software across multi-GPU and multi-node training stacks, enabling GenAI/LLM scaling with high performance and reliability.

You will lead technical initiatives, collaborate with ML engineers and infrastructure teams, and contribute to performance benchmarks and optimizations across CUDA and PyTorch

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Distributed ML Systems Engineer - Scaling (Equity)
Distributed ML Systems Engineer - Scaling (Equity)

Meta • Menlo Park (CA)

On-site
USD 154,000 - 217,000
Bonus
Equity
Benefits
Software Engineer, SystemML - AI Networking
Software Engineer, SystemML - AI Networking

Meta • Menlo Park (CA)

On-site
USD 154,000 - 217,000
Bonus
Equity
Benefits
Software Engineer, SystemML - Scaling / Performance
Software Engineer, SystemML - Scaling / Performance

Meta • Menlo Park (CA)

On-site
USD 154,000 - 217,000
Bonus
Equity
Benefits
Research Scientist, AI Networking (PhD)
Research Scientist, AI Networking (PhD)

Meta • Menlo Park (CA)

On-site
USD 121,000 - 181,000
Senior AI Networking & Performance Engineer
Senior AI Networking & Performance Engineer

NVIDIA • Town of Texas (WI)

On-site
USD 272,000 - 432,000
Senior AI Networking and Performance Engineer for LLMs
Senior AI Networking and Performance Engineer for LLMs

Nvidia Corporation • Santa Clara (CA)

On-site
USD 272,000 - 431,250
Equity
Comprehensive benefits
AI/HPC Performance Engineer: Scale Large AI Clusters
AI/HPC Performance Engineer: Scale Large AI Clusters

Meta • Menlo Park (CA)

On-site
USD 154,000 - 217,000
Bonus
Equity
Benefits
ML Systems Engineer: AI Infra & GPU Acceleration
ML Systems Engineer: AI Infra & GPU Acceleration

Meta • San Francisco (CA)

On-site
USD 180,000 - 240,000
Bonus
Equity
AI/ML Network Infrastructure Engineer I
AI/ML Network Infrastructure Engineer I

Amazon • Cupertino (CA)

On-site
USD 127,000 - 185,000
Senior AI Networking Software Engineer (Remote)
Senior AI Networking Software Engineer (Remote)

Cornelis Networks • United States

On-site
USD 150,000 - 230,000
Equity
Health benefits
401(k) with company match
+1