Research Engineer, Large-Scale Training

Together AI

San Francisco (CA)

On-site

USD 200,000 - 290,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health insurance
Startup equity
Other benefits

Job summary

Together AI is seeking a Research Engineer on the Scaling Team to translate cutting-edge research into robust training systems. You will profile and optimize training infrastructure, identify bottlenecks, and implement techniques from literature and our scientists into production.

You will enable support for open-source foundation models on the Together platform, collaborate with researchers to accelerate experiments, and help deploy validated ideas at scale.

Qualifications

  • Ability to independently take ambiguous performance problems from investigation through deployment.
  • Strong Python and PyTorch skills, with efficient, maintainable code.
  • Hands-on experience training or fine-tuning large neural networks in multi-GPU or multi-node setups.

Responsibilities

  • Design, implement, and optimize core components of Together's large-scale training infrastructure.
  • Integrate new model architectures, validate training correctness and convergence, and optimize performance for production fine-tuning workloads.
  • Profile distributed training workloads to identify and eliminate bottlenecks across compute, memory, and communication.
  • Design and execute experiments to validate performance hypotheses and benchmark new approaches against state-of-the-art methods.
  • Partner closely with Research Scientists to productionize novel training methods and contribute to publications and open-source releases.
  • Rapidly enable support for newly released open-source foundation models on the Together platform.
  • Build and maintain experimental infrastructure that accelerates research while ensuring production-quality reliability and scalability.

Skills

Python
PyTorch
Multi-GPU
ML systems
Communication
Productionizing

Tools

CUDA
NCCL
DeepSpeed
Megatron-LM
FSDP

Job description

The Model Shaping team at Together AI works on products and research for tailoring open foundation models to downstream applications. We build services that allow machine learning developers to choose the best models for their tasks and further improve these models using domain-specific data. In addition, we develop new methods for more efficient model training and evaluation, drawing inspiration from a broad spectrum of ideas across machine learning, natural language processing, and ML systems.

As a Research Engineer on the Scaling Team within Model Shaping, you will turn cutting-edge research on efficient foundation model training into robust, high-performance systems. You will profile and optimize Together's training infrastructure, identify performance bottlenecks across the stack, and implement state-of-the-art techniques from both the research literature and our own scientists in production environments.

Your work will directly shape the fine-tuning experience of Together's customers. You will rapidly bring newly released open-source models onto the Model Shaping platform, ensuring they train efficiently and reliably across diverse customer workloads. Working closely with Research Scientists, you will also build the experimental infrastructure that accelerates research and enables validated ideas to be deployed reliably at scale.

Responsibilities
  • Design, implement, and optimize core components of Together's large-scale training infrastructure.
  • Integrate new model architectures, validate training correctness and convergence, and optimize performance for production fine-tuning workloads.
  • Profile distributed training workloads to identify and eliminate bottlenecks across compute, memory, and communication.
  • Design and execute experiments to validate performance hypotheses and benchmark new approaches against state-of-the-art methods.
  • Partner closely with Research Scientists to productionize novel training methods and contribute to publications and open-source releases.
  • Rapidly enable support for newly released open-source foundation models on the Together platform.
  • Build and maintain experimental infrastructure that accelerates research while ensuring production-quality reliability and scalability.
Requirements
  • Demonstrated ability to independently take ambiguous performance or infrastructure problems from investigation through deployment.
  • Strong programming skills in Python and PyTorch, with an emphasis on writing efficient, maintainable code.
  • Hands‑on experience training or fine‑tuning large neural networks in multi‑GPU or multi‑node environments.
  • Solid understanding of ML systems fundamentals, including GPU architecture, mixed‑precision training, and distributed training paradigms such as data, tensor, pipeline, or expert parallelism.
  • Strong communication skills and the ability to collaborate effectively with both researchers and engineers.
  • Passion for staying current with advances in AI research and applying them to real‑world systems.
  • Excitement about translating cutting‑edge research into production systems that deliver customer impact.
Nice to Have
  • Experience writing optimized NVIDIA GPU kernels using CUDA or Triton, or implementing communication collectives with technologies such as NCCL or NVSHMEM.
  • Experience with large‑scale training frameworks such as FSDP, DeepSpeed, Megatron‑LM, or custom distributed training systems.
  • Experience optimizing distributed training for compute efficiency, memory efficiency, or scalability.
  • Experience running and managing large-scale GPU experiments, including scheduling, monitoring, and fault tolerance.
  • Contributions to widely used open‑source ML or ML systems projects.
  • Experience building or operating ML products or managed services used by external customers.
About Together AI

Together AI is a research‑driven artificial intelligence company. We believe open and transparent AI systems will drive innovation and create the best outcomes for society, and together we are on a mission to significantly lower the cost of modern AI systems by co‑designing software, hardware, algorithms, and models. We have contributed to leading open‑source research, models, and datasets to advance the frontier of AI, and our team has been behind technological advancement such as FlashAttention, ATLAS, RedPajama, and Mamba. We invite you to join a passionate group of researchers in our journey in building the next generation AI infrastructure.

Compensation

We offer competitive compensation, startup equity, health insurance, and other benefits. The US base salary range for this full‑time position is $200,000 - $290,000. Our salary ranges are determined by location, level and role. Individual compensation will be determined by experience, skills, and job‑related knowledge.

Together AI is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more.

Are you willing to work four days per week in our San Francisco office?

Are you legally authorized to work in the country where the job is located?

Will you now or in the future require company sponsorship to retain or extend your work authorization in the country where the job is located?

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Research Engineer, Large-Scale Training
Research Engineer, Large-Scale Training

Togetherai • San Francisco (CA)

On-site
USD 200,000 - 290,000
Competitive compensation
Startup equity
Health insurance
+1
Research Engineer, Large-Scale Training Together AI San Francisco $200,000 - $290,000/yr
Research Engineer, Large-Scale Training Together AI San Francisco $200,000 - $290,000/yr

Neura Market • San Francisco (CA), Northern (KY)

Hybrid
USD 200,000 - 290,000
Startup equity
Health insurance
Other benefits
Research Engineer, Post-Training Inference
Research Engineer, Post-Training Inference

Togetherai • San Francisco (CA)

On-site
USD 200,000 - 290,000
Startup equity
Health insurance
Research Engineer, Post-Training Inference
Research Engineer, Post-Training Inference

Together AI • San Francisco (CA)

On-site
USD 200,000 - 290,000
Health insurance
Startup equity
Platform Engineer, Model Shaping
Platform Engineer, Model Shaping

Togetherai • San Francisco (CA)

Hybrid
USD 200,000 - 290,000
Competitive compensation
Startup equity
Health insurance
+1
Platform Engineer, Model Shaping
Platform Engineer, Model Shaping

Together AI • San Francisco (CA)

Hybrid
USD 200,000 - 290,000
Startup equity
Health insurance
Flexible remote work
AI Researcher, Core ML (Turbo)
AI Researcher, Core ML (Turbo)

Togetherai • San Francisco (CA)

On-site
USD 200,000 - 280,000
Startup equity
Health insurance
Competitive benefits
Research Engineer, Core ML
Research Engineer, Core ML

Togetherai • San Francisco (CA)

On-site
USD 200,000 - 280,000
Startup equity
Health insurance
Competitive benefits
AI Researcher, Core ML (Turbo)
AI Researcher, Core ML (Turbo)

Together • San Francisco (CA)

On-site
USD 200,000 - 280,000
Health insurance
Startup equity
Competitive benefits
Research Intern, Model Shaping (Fall 2026)
Research Intern, Model Shaping (Fall 2026)

Together AI • San Francisco (CA)

Hybrid
Competitive compensation
Housing stipends
Other competitive benefits