AI Training Optimization Scientist: Scale Efficiently

United States Digital Space LLC

United States

Remote

USD 140,000 - 230,000

Full time

2 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

United States Digital Space LLC is seeking an AI Researcher focused on training optimization to push efficiency and scalability of large-scale model training. You will work at the intersection of research and systems to develop techniques to reduce training cost and accelerate convergence.

Ideal candidates have a track record (or strong ambition) of publishing applied ML research and will collaborate with infrastructure and inference teams to translate research into real-world performance.

Qualifications

  • Strong background in ML research with emphasis on training dynamics and optimization.
  • Experience training large neural networks (LLMs, multimodal models, or long sequence models).
  • Publication experience in ML venues or equivalent high-quality open research.
  • Proficiency in Python and modern ML frameworks (PyTorch preferred).

Responsibilities

  • Design and evaluate training optimization techniques for large models (e.g. optimization algorithms, schedulers, normalization, curriculum strategies).
  • Run large-scale experiments, analyze results, and translate findings into actionable improvements.
  • Collaborate with infrastructure and inference teams to ensure training decisions translate to real-world performance.
  • Author or co-author research papers, technical reports, or blog posts.

Skills

Training optimization
Large models
Python
Experiment design
Distributed training
Publications

Tools

PyTorch

Job description

United States Digital Space LLC is seeking an AI Researcher focused on training optimization to push efficiency and scalability of large-scale model training. You will work at the intersection of research and systems to develop techniques to reduce training cost and accelerate convergence.

Ideal candidates have a track record (or strong ambition) of publishing applied ML research and will collaborate with infrastructure and inference teams to translate research into real-world performance.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Training Infrastructure Engineer
AI Training Infrastructure Engineer

United States Digital Space LLC • United States

Remote
USD 150,000 - 190,000
AI Training Infrastructure Engineer - Scale & Performance
AI Training Infrastructure Engineer - Scale & Performance

SupportFinity™ • San Francisco (CA)

On-site
USD 175,000 - 220,000
Meaningful equity
Competitive salary
Comprehensive benefits package
Senior ML Engineer - Robotics Autonomy & Large-Scale Training
Senior ML Engineer - Robotics Autonomy & Large-Scale Training

United States Digital Space LLC • United States

Remote
USD 150,000 - 210,000
AI Training Infrastructure Engineer - Scale LLM Training
AI Training Infrastructure Engineer - Scale LLM Training

Fireworks AI • San Mateo (CA)

On-site
USD 175,000 - 220,000
Equity
Comprehensive benefits
Competitive salary
Senior AI Training Performance Architect-Optimization Lead
Senior AI Training Performance Architect-Optimization Lead

Nvidia Corporation • Santa Clara (CA)

On-site
USD 184,000 - 356,500
AI Performance Engineer: Scale ML Training & Inference
AI Performance Engineer: Scale ML Training & Inference

applied • Sunnyvale (CA)

On-site
USD 180,000 - 240,000
Tech Lead Manager for Scalable LLM Training Platform
Tech Lead Manager for Scalable LLM Training Platform

United States Digital Space LLC • San Francisco (CA), New York (NY)

On-site
USD 290,000 - 363,000
AI Infrastructure Engineer: Scale Training & Systems
AI Infrastructure Engineer: Scale Training & Systems

Precision Labs • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
LLM Systems Engineer & Research
LLM Systems Engineer & Research

Scale AI, Inc. • New York (NY)

On-site
USD 189,000 - 237,000
Comprehensive health coverage
Dental and vision coverage
Retirement benefits
+3
Remote GPU Performance Engineer for Large-Scale Models
Remote GPU Performance Engineer for Large-Scale Models

United States Digital Space LLC • United States

Remote
USD 140,000 - 210,000
Five weeks paid leave
Comprehensive healthcare (vision +</p>