Pre-training Research Engineer

Sciforium

San Francisco (CA)

On-site

USD 180,000 - 260,000

Full time

11 days ago
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Medical, dental, and vision insurance
401k plan
Daily lunch, snacks, and beverages
Flexible time off
Competitive salary and equity

Job summary

Sciforium in San Francisco is seeking a Pre-training Research Engineer to advance byte-native and multimodal foundation models. You will implement training code, scale experiments, and contribute production-grade infrastructure to deliver high-quality base models for real-world use cases.

The role emphasizes rapid iteration, rigorous ablations, and collaboration with data and distributed training engineers to improve efficiency, reliability, and scalability across large GPU clusters.

Qualifications

  • 5+ years of experience in ML research or engineering with a track record of pre-training foundation models.
  • Strong software engineering skills to write robust training code.
  • Solid understanding of deep learning fundamentals and pre-training methods.
  • Ability to implement research ideas and evaluate with baselines and metrics.
  • Hands-on experience with GPU-based training and distributed training workflows.

Responsibilities

  • Train large byte-native and multimodal foundation models across massive corpora.
  • Implement and evaluate new model architectures, training objectives, and optimizations.
  • Develop stable pre-training recipes and run scaling experiments for novel architectures.
  • Conduct ablations and analyze training dynamics and base-model quality.
  • Work with data and distributed training engineers to improve training efficiency and scalability.

Skills

5+ years ML exp
Software engineering
ML Foundations
Research & experimentation
GPU & distributed training

Education

MS in CS/ML/AI/Math

Tools

JAX
Flax
XLA
FSDP
Megatron

Job description

Sciforium is an AI infrastructure company developing next-generation multimodal AI models and a proprietary, high-efficiency serving platform. Backed by multi-million-dollar funding and direct sponsorship from AMD with hands‑on support from AMD engineers the team is scaling rapidly to build the full stack powering frontier AI models and real-time applications.

About the role

As a Pre-training Research Engineer, you’ll focus on model implementation, pertaining and scaling, and improving the quality of our byte-native and multimodal foundation models. You’ll build and iterate quickly on research ideas, contribute production‑grade training code and infrastructure, and help deliver high‑quality base models that can serve real‑world use cases at scale.

Key Responsibilities
Pre-training & Scaling
  • Train large byte-native and multimodal foundation models across massive, heterogeneous corpora.

  • Implement and evaluate new model architectures, training objectives, and optimization methods.

  • Develop stable pre‑training recipes and run scaling experiments for novel architectures.

  • Conduct ablations and analyze training dynamics, model behavior, and base-model quality.

  • Work with data and distributed training engineers to improve training efficiency, reliability, and scalability.

Must-Haves
  • 5+ years of experience in machine learning research or engineering, with a proven track record of developing and pre‑training large language or multimodal foundation models.

  • Software Engineering: Strong general software engineering skills, with the ability to write robust and performant training code.

  • ML Foundations: Solid understanding of deep learning fundamentals and modern pre‑training methods and literature.

  • Research and Experimentation: Ability to quickly implement research ideas and evaluate them using clear baselines, ablations, metrics, and analysis.

  • GPU and Distributed Training: Hands‑on experience running training workloads in GPU‑based environments, with familiarity with distributed training.

  • Education: MS in Computer Science, Machine Learning, Artificial Intelligence, Mathematics, or a related field.

Nice-to-Haves
  • PhD in Computer Science, Machine Learning, Artificial Intelligence, Mathematics, or a related field.

  • JAX Ecosystem: Extensive experience with the JAX, Flax, and XLA stack.

  • Large‑Scale Distributed Training: Experience with multi‑node pre‑training using systems such as FSDP, ZeRO, or Megatron.

  • Training Recipes and Scaling: Experience developing training recipes, ablations, or scaling experiments.

  • Monitoring and Reproducibility: Experience owning end‑to‑end training and evaluation pipelines with monitoring and reproducibility.

Education
  • MS or PhD in Computer Science, Machine Learning, Artificial Intelligence, Mathematics, or a related field.

Benefits include
  • Medical, dental, and vision insurance

  • 401k plan

  • Daily lunch, snacks, and beverages

  • Flexible time off

  • Competitive salary and equity

Equal opportunity

Sciforium is an equal opportunity employer. All applicants will be considered for employment without attention to race, color, religion, sex, sexual orientation, gender identity, national origin, veteran or disability status.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Research Scientist
Senior Research Scientist

Sciforium • San Francisco (CA)

On-site
USD 130,000 - 160,000
Medical, dental, and vision insurance
401k plan
Daily lunch, snacks, and beverages
+2
Lead Software Engineer, Model Serving Platform
Lead Software Engineer, Model Serving Platform

Sciforium • San Francisco (CA)

On-site
USD 180,000 - 240,000
Medical, dental, and vision insurance
401k plan
Daily lunch, snacks, and beverages
+2
Research Engineer - Model Evaluation & MLOps
Research Engineer - Model Evaluation & MLOps

Sciforium • San Francisco (CA)

On-site
USD 140,000 - 200,000
Medical, dental, and vision insurance
401k plan
Daily lunch, snacks, and beverages
+2
Senior AI Serving Engineer, Backend
Senior AI Serving Engineer, Backend

Sciforium • San Francisco (CA)

On-site
USD 190,000 - 250,000
Medical, dental, and vision insurance
401(k) plan
Daily lunch, snacks, and beverages
+2
Senior Pre-training Research Engineer - Foundation Models
Senior Pre-training Research Engineer - Foundation Models

Sciforium • San Francisco (CA)

On-site
USD 180,000 - 260,000
Medical, dental, and vision insurance
401k plan
Daily lunch, snacks, and beverages
+2
ML Platform Engineer, Backend
ML Platform Engineer, Backend

Sciforium • San Francisco (CA)

On-site
USD 170,000 - 230,000
Medical, dental, and vision insurance
401k plan
Daily lunch, snacks, and beverages
+2
Software Engineer, Fullstack
Software Engineer, Fullstack

Sciforium • San Francisco (CA)

On-site
USD 165,000 - 210,000
Medical, dental, and vision insurance
401k plan
Daily lunch, snacks, and beverages
+2
Research Scientist - Distributed Machine Learning
Research Scientist - Distributed Machine Learning

Institute of Foundation Models • Sunnyvale (CA)

On-site
USD 300,000 - 600,000
Comprehensive medical, dental, and vision benefits
401K Plan
Generous paid time off
+3
Foundation Model Data Engineer
Foundation Model Data Engineer

Sciforium • San Francisco (CA)

On-site
USD 155,000 - 210,000
Medical, dental, and vision insurance
401k plan
Daily lunch, snacks, and beverages
+2
Research Engineer - Midtraining
Research Engineer - Midtraining

Periodic Labs • Menlo Park (CA)

On-site
USD 250,000 - 350,000