Distributed ML Systems Engineer - Scaling (Equity)

Meta

Menlo Park (CA)

On-site

USD 154,000 - 217,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Bonus
Equity
Benefits

Job summary

Meta's Network.AI Software team seeks a Software Engineer, SystemML – Scaling / Performance to enable reliable, scalable distributed ML training on Meta's large-scale GPU training infrastructure, focusing on GenAI/LLM scaling across inter-GPU and network layers.

You will work with NCCL-based stacks, contribute to performance benchmarks and optimizations, and collaborate with PyTorch and ML framework teams to improve full-stack reliability and throughput for distributed AI workloads.

Qualifications

  • Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience.
  • Specialized experience in one or more of: Distributed ML Training, GPU architecture, ML systems, AI infrastructure, high performance computing, performance optimizations, or ML frameworks (e.g. PyTorch).

Responsibilities

  • Enabling reliable and highly scalable distributed ML training on Meta's large-scale GPU training infra with a focus on GenAI/LLM scaling

Skills

Distributed ML Training
GPU architecture
ML systems
AI infrastructure
High performance computing
Performance optimizations
ML frameworks (e.g. PyTorch)

Education

Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience

Tools

CUDA programming
PyTorch
Caffe2
TensorFlow

Job description

Meta's Network.AI Software team seeks a Software Engineer, SystemML – Scaling / Performance to enable reliable, scalable distributed ML training on Meta's large-scale GPU training infrastructure, focusing on GenAI/LLM scaling across inter-GPU and network layers.

You will work with NCCL-based stacks, contribute to performance benchmarks and optimizations, and collaborate with PyTorch and ML framework teams to improve full-stack reliability and throughput for distributed AI workloads.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineer - AI Networking & Distributed GPU Systems
Software Engineer - AI Networking & Distributed GPU Systems

Meta • Menlo Park (CA)

On-site
USD 154,000 - 217,000
Bonus
Equity
Benefits
Staff ML Systems Engineer - Scalable AI Infrastructure
Staff ML Systems Engineer - Scalable AI Infrastructure

Meta • Menlo Park (CA)

On-site
USD 183,000 - 257,000
Software Engineer, SystemML - Scaling / Performance
Software Engineer, SystemML - Scaling / Performance

Meta • Menlo Park (CA)

On-site
USD 154,000 - 217,000
Bonus
Equity
Benefits
Senior ML Systems Engineer — Scalable AI Infra
Senior ML Systems Engineer — Scalable AI Infra

Meta • Menlo Park (CA)

On-site
USD 347,000 - 403,000
Senior ML Engineer: Scalable AI for Global Impact
Senior ML Engineer: Scalable AI for Global Impact

Jobzhr • New York (NY)

On-site
USD 184,000 - 257,000
Bonus
Equity
Benefits
ML Systems Engineer: AI Infra & GPU Acceleration
ML Systems Engineer: AI Infra & GPU Acceleration

Meta • San Francisco (CA)

On-site
USD 180,000 - 240,000
Bonus
Equity
Software Engineer, SystemML - AI Networking
Software Engineer, SystemML - AI Networking

Meta • Menlo Park (CA)

On-site
USD 154,000 - 217,000
Bonus
Equity
Benefits
AI/HPC Performance Engineer: Scale Large AI Clusters
AI/HPC Performance Engineer: Scale Large AI Clusters

Meta • Menlo Park (CA)

On-site
USD 154,000 - 217,000
Bonus
Equity
Benefits
Senior AI Infrastructure Engineer
Senior AI Infrastructure Engineer

AI Breaking Wire • Menlo Park (CA), Northern (KY)

Hybrid
USD 200,000 - 350,000
RSUs
Health benefits
Parental leave
+1
Senior ML Infra Engineer - Scale GPU Clusters, Remote
Senior ML Infra Engineer - Scale GPU Clusters, Remote

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 320,000 - 500,000
Equity
Medical/Dental/Vision coverage
Unlimited PTO
+1