AI Platform Engineer: Scale Large Models

LinkedIn

Mountain View (CA)

Hybrid

USD 120,000 - 195,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

LinkedIn is hiring for an AI Training Infra engineer to scale training, feature engineering, and serving for billions‑parameter models. You will design high‑performance data I/O, work with PyTorch and open source libraries, enable distributed training across hundreds of billions of parameters, and optimize GPU and container infrastructure.

The role emphasizes collaboration with researchers and engineers to advance one of the world’s most scalable AI platforms, while advancing your expertise in

Qualifications

  • Bachelor’s degree in Computer Science or a related technical discipline or equivalent practical experience.
  • 1+ years in the industry with leading/building deep learning systems.
  • Experience with Java, C++, Python, Go, Rust, C# and/or Scala or other relevant coding languages.
  • Experience, qualifications in Machine Learning, AI.

Responsibilities

  • Design, implement, and optimize the performance of large‑scale distributed serving or training for personalized recommendation as well as large language models.
  • Improving the observability and understandability of various systems with a focus on improving developer productivity and system sustenance.
  • Partner with peers, leads and partners to define, scope, prioritize, and build impactful features at a high velocity.

Skills

Distributed systems
Java
Python
Go
Rust
C#
Scala

Education

Bachelor's degree in CS or related field
MS/PhD in CS or related technical discipline

Tools

PyTorch
TensorFlow
HuggingFace
Horovod
DeepSpeed
PyTorch Lightning
CUDA
cuDNN
NCCL
Kubernetes
Spark
Beam
Flink

Job description

LinkedIn is hiring for an AI Training Infra engineer to scale training, feature engineering, and serving for billions‑parameter models. You will design high‑performance data I/O, work with PyTorch and open source libraries, enable distributed training across hundreds of billions of parameters, and optimize GPU and container infrastructure.

The role emphasizes collaboration with researchers and engineers to advance one of the world’s most scalable AI platforms, while advancing your expertise in

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Platform Engineer - Scalable ML Infra
AI Platform Engineer - Scalable ML Infra

LinkedIn • Mountain View (CA)

Hybrid
USD 120,000 - 195,000
Senior ML Engineer — Scale Training Infra & AI Deployments
Senior ML Engineer — Scale Training Infra & AI Deployments

Best AI Tools Wiki • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 350,000
Equity package
Health insurance
Unlimited PTO
+3
AI Infra Engineer: Build Scalable Model Serving
AI Infra Engineer: Build Scalable Model Serving

Deep Infra Inc. • Palo Alto (CA)

On-site
USD 150,000 - 195,000
Open-source contributions
C++/CUDA experience
Staff AI Platform Engineer: Scale ML Infra
Staff AI Platform Engineer: Scale ML Infra

DAT Freight Solutions • Seattle (WA)

Hybrid
USD 198,000 - 246,000
Medical Insurance
Dental Insurance
Vision Insurance
+6
AI Infrastructure Engineer — Scale ML Training & Inference
AI Infrastructure Engineer — Scale ML Training & Inference

Triwill Group • San Francisco (CA), Northern (KY)

Hybrid
USD 140,000 - 190,000
AI Infrastructure Engineer: Scale Training & Systems
AI Infrastructure Engineer: Scale Training & Systems

Precision Labs • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Scale & Infrastructure Engineer for Large-Scale AI
Scale & Infrastructure Engineer for Large-Scale AI

AI Breaking Wire • Mountain View (CA), Northern (KY)

Hybrid
USD 165,000 - 230,000
Stock options
Healthcare coverage
PTO
+1
Platform Engineer: Model Shaping & Scalable AI Infra
Platform Engineer: Model Shaping & Scalable AI Infra

Together AI • San Francisco (CA)

On-site
USD 120,000 - 150,000
Competitive health insurance plans
401(k) plan
Flexible time off policy
+1
AI Data Platform Engineer: Scalable Systems
AI Data Platform Engineer: Scalable Systems

Amazon • Palo Alto (CA)

On-site
USD 165,000 - 224,000
Senior AI Infrastructure Architect for Scalable Training
Senior AI Infrastructure Architect for Scalable Training

LinkedIn • Sunnyvale (CA)

Hybrid
USD 198,000 - 326,000