Member of Technical Staff, Applied Research

Sieve

San Francisco (CA)

On-site

USD 180,000 - 280,000

Full time

6 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

401k + Full Health Insurance
Breakfast, Lunch, and Dinner covered
Ubers covered home

Job summary

Sieve is seeking a Member of Technical Staff, Applied Research, in San Francisco to train and evaluate multimodal models. You will connect data curation with model performance, own the research loop from hypothesis to evaluation, and build reproducible training pipelines for video generation and multimodal tasks.

You will bridge research and engineering, implement methods from papers, debug training runs, and drive reliable systems for sourcing, curating, and improving training data.

Qualifications

  • 2+ years of experience in machine learning research or engineering.
  • Strong Python and PyTorch skills, including the ability to implement, debug, and modify model training code.
  • Experience designing experiments, establishing baselines, and evaluating results critically.
  • Familiarity with modern generative or multimodal architectures, such as diffusion models or transformers.
  • Comfortable working with large datasets and diagnosing training bottlenecks, instability, and data quality issues.
  • Able to turn ambiguous research questions into concrete experiments and maintainable systems.
  • Strong communication skills and the ability to explain findings, tradeoffs, and uncertainty clearly.

Responsibilities

  • Train and post-train models for video generation and multimodal understanding.
  • Design controlled experiments to measure how data selection, mixtures, and supervision affect model capabilities.
  • Build evaluations that reveal specific model weaknesses, using quantitative metrics and human judgment.
  • Develop reliable training infrastructure, including distributed training, efficient data loading, checkpointing, and experiment tracking.
  • Turn research findings into improvements in data curation pipelines and products.
  • Collaborate with research and engineering teams internally and at partner labs to define meaningful problems and communicate results.

Skills

Python
PyTorch
Experiment design
Diffusion models
Multimodal ML
Large datasets
Communication

Job description

About Us

Sieve is a multi-modal lab curating the world's highest-quality training datasets — spanning video, audio, images, text, and 3D. We combine exabyte-scale data infrastructure and novel multimodal understanding techniques that push the frontier of foundation models. Video alone makes up 80% of internet traffic, and across modalities, data has become the enabling medium powering creativity, communication, gaming, AR/VR, and robotics. Sieve exists to solve the biggest bottleneck in the growth of these applications: high-quality training data.

We partner with top AI labs and did $XXM last quarter alone, as a team of ~30 people. We also raised our Series A from Tier 1 firms such as Matrix Partners, Swift Ventures, Y Combinator, and AI Grant.

Why Now

Sieve combines access to diverse multimodal data, infrastructure to process it at scale, and close relationships with the teams building frontier models. This gives us a unique opportunity to study what makes training data effective—and turn those findings into better models and datasets.

You’ll join a small team building our research capabilities, with ownership over experiments, training systems, and the decisions those results inform.

About the Role

As a Member of Technical Staff, Applied Research at Sieve, you’ll train and evaluate multimodal models to understand how data shapes their capabilities. Your work will span video generation and audiovisual understanding, connecting advances in data curation with measurable improvements in model performance.

You’ll own the research loop end-to-end: identify a model weakness, form a hypothesis about the data or training approach that could address it, build the experiment, and evaluate the results. This includes fine-tuning and post-training models, developing reproducible training and evaluation pipelines, and running controlled experiments on data quality, composition, and supervision.

You’re likely a good fit if you enjoy moving between research and engineering: reading a paper, implementing a method, debugging a training run, and figuring out whether an apparent improvement holds up. You care about building reliable systems and producing findings that change how we source, curate, and use training data.

What You’ll Work On

  • Train and post-train models for video generation and multimodal understanding.
  • Design controlled experiments to measure how data selection, mixtures, and supervision affect model capabilities.
  • Build evaluations that reveal specific model weaknesses, using quantitative metrics and human judgment.
  • Develop reliable training infrastructure, including distributed training, efficient data loading, checkpointing, and experiment tracking.
  • Turn research findings into improvements in our data curation pipelines and products.
  • Collaborate with research and engineering teams internally and at partner labs to define meaningful problems and communicate results.
Requirements
  • 2+ years of experience in machine learning research or engineering, with hands-on experience training or fine-tuning deep learning models.
  • Strong Python and PyTorch skills, including the ability to implement, debug, and modify model training code.
  • Experience designing experiments, establishing baselines, and evaluating results critically.
  • Familiarity with modern generative or multimodal architectures, such as diffusion models or transformers.
  • Comfortable working with large datasets and diagnosing training bottlenecks, instability, and data quality issues.
  • Able to turn ambiguous research questions into concrete experiments and maintainable systems.
  • Strong communication skills and the ability to explain findings, tradeoffs, and uncertainty clearly.
Bonus
  • Experience training video, image, or audio generation models.
  • Experience with supervised fine-tuning, preference optimization, or reinforcement learning.
  • Experience with distributed training and GPU performance optimization.
  • Research publications, open-source contributions, or substantial independent ML projects.
  • Experience as an early hire at a startup.
Benefits
  • 401k + Full Health Insurance
  • Breakfast, Lunch, and Dinner covered and your choice of snacks
  • Ubers covered home

*all roles at Sieve require you to be onsite in San Francisco 5 days per week

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Member of Technical Staff, Machine Learning
Member of Technical Staff, Machine Learning

Sieve, Inc. • San Francisco (CA)

On-site
USD 150,000 - 230,000
Member of Technical Staff, Machine Learning
Member of Technical Staff, Machine Learning

Sieve • San Francisco (CA)

On-site
USD 150,000 - 190,000
Member of Technical Staff, Infrastructure
Member of Technical Staff, Infrastructure

Sieve • San Francisco (CA)

On-site
USD 150,000 - 210,000
401k
Health insurance
Meals & snacks
+2
Member of Technical Staff, Deployed Research
Member of Technical Staff, Deployed Research

Sieve, Inc. • San Francisco (CA)

On-site
USD 100,000 - 140,000
401k + Full Health Insurance
Meals covered (Breakfast, Lunch, and Dinner)
Uber rides covered home
Member of Technical Staff
Member of Technical Staff

Sieve • San Francisco (CA)

On-site
USD 150,000 - 350,000
401k + Full Health Insurance
Meals covered
Ubers covered home
Product & Ops Lead
Product & Ops Lead

Sieve • San Francisco (CA)

On-site
USD 90,000 - 130,000
401k + Full Health Insurance
Meals covered including snacks
Uber rides home covered
Member of Technical Staff, Forward Deployed
Member of Technical Staff, Forward Deployed

Sieve • San Francisco (CA)

On-site
USD 100,000 - 140,000
401k + Full Health Insurance
Meals covered
Uber rides covered home
Business Development
Business Development

Sieve, Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 120,000 - 180,000
401(k) and full health insurance
Breakfast, lunch, and dinner provided
Onsite in San Francisco 5 days a week
Business Development
Business Development

Sieve • San Francisco (CA)

On-site
USD 175,000 - 275,000
401(k) and full health insurance
Meals provided (breakfast, lunch, and
Ubers covered
+1
Research & Product Lead
Research & Product Lead

Sieve • San Francisco (CA)

On-site
USD 120,000 - 160,000
401k + Full Health Insurance
Meals provided
Uber rides covered home