Research Scientist

Stealth Startup

San Francisco (CA)

On-site

USD 180,000 - 280,000

Full time

44 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Stealth Startup in San Francisco seeks a Research Scientist to advance foundation models for tabular data, time series, and broader structured data. You will own the research agenda, design experiments, and write up results for engineers and stakeholders.

You will drive end-to-end pretraining, develop synthetic data strategies, and build robust evaluation systems while staying current with cutting-edge research. This is a high-ownership, high-freedom role in an early-stage environment.

Qualifications

  • A strong publication record; PhD preferred but not required.
  • Hands-on experience training large-scale models with multi-GPU or multi-node pretraining runs.
  • Comfortable reading and modifying attention kernels, developing custom modules, and profiling systems.
  • Deep familiarity with transformer internals and architectural trade-offs.
  • Strong written and verbal communication skills with cross-functional colleagues.
  • Comfortable in an early-stage environment where priorities evolve.

Responsibilities

  • Design and run experiments on model architecture, including attention mechanisms and in-context learning for tabular foundation models.
  • Own pretraining runs end to end—launch, profile, debug loss curves, and analyze results.
  • Design synthetic data generation strategies to control generalization capabilities.
  • Build evaluation systems, benchmarks, scoring rules, and perform statistical analyses.
  • Extend models to new modalities and translate research into concrete decisions.

Skills

Research
Publications
Communication
Problem solving
Transformer knowledge

Education

PhD preferred

Tools

PyTorch
CUDA
Distributed training
FlashAttention

Job description

We are looking for a Research Scientist to help build next-generation foundation models for tabular data, time series, and ultimately the broader world of structured data.

Our goal is to make powerful machine learning dramatically easier to use, customize, and deploy across real-world applications and enterprise environments.

This is a high-ownership, high-freedom research role. You'll drive your own research agenda across model architecture, pretraining, synthetic data, and evaluation. You'll own training runs end to end- from the initial idea and ablation studies to analysis and technical writeups. Just as importantly, you'll learn alongside your colleagues: a negative result is information, not a failure.

Responsibilities
  • Design and run experiments on model architecture, including attention mechanisms, normalization, positional schemes, optimizers, and the in-context learning machinery that makes tabular foundation models work.
  • Own pretraining runs end to end-launch them, profile them, debug loss curves, and understand why the model behaved the way it did.
  • Design synthetic data generation strategies, including priors, sampling, and reweighting choices that determine what the model can and cannot generalize to.
  • Build evaluation systems that measure meaningful performance. Design benchmarks, select appropriate scoring rules, run statistical analyses rigorously, and identify when a result is unexpectedly strong or potentially misleading.
  • Extend models to new modalities.
  • Stay closely engaged with current research and translate it into concrete decisions: determine which recent techniques are worth testing in our environment and which results may not translate beyond the original paper.
  • Write up research findings through internal reports, white papers, and technical blog posts, communicating clearly with engineers, customers, and reviewers.
  • Help define the research roadmap and make foundational modeling decisions.
Qualifications
  • A strong publication record in machine learning or an equivalent public body of work, such as open-source models, technical reports, or reproducible results that have gained recognition. A PhD is preferred but not required.
  • Hands-on experience training large-scale models. You have owned multi‑GPU or multi-node pretraining runs and understand the stability, throughput, and data‑pipeline challenges involved.
  • Comfortable reading and modifying attention kernels, developing custom modules, and profiling systems to identify performance bottlenecks.
  • Deep familiarity with transformer internals and the trade‑offs associated with different architectural choices.
  • Strong written and verbal communication skills. You can explain technical mechanisms precisely without relying on vague explanations and can communicate effectively with both technical and cross‑functional colleagues.
  • Comfortable working in an early‑stage environment where priorities evolve and requirements are not always fully defined.
Nice to Have
  • Experience with tabular foundation models or in‑context learning.
  • Experience with time‑series forecasting or probabilistic prediction.
  • GPU performance optimization experience, including FlashAttention‑style kernels, sparse or linear attention, memory optimization, and distributed training using technologies such as FSDP and torchrun.
  • Experience with synthetic data generation for pretraining.
  • A track record of open-source contributions or maintaining research code.
  • Previous experience at an early‑stage startup or as part of a founding team.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Research Scientist: Pretraining
Research Scientist: Pretraining

Generalist AI • San Mateo (CA), Somerville (MA)

On-site
USD 180,000 - 240,000
Research Scientist: Foundational Models for Structured Data
Research Scientist: Foundational Models for Structured Data

Stealth Startup • San Francisco (CA)

On-site
USD 180,000 - 280,000
Research Scientist Intern (PhD)
Research Scientist Intern (PhD)

Prior Labs • New York (NY)

On-site
USD 120,000 - 180,000
Mentorship and professional growth
State-of-the-art ML architecture and <
Healthcare, transportation, and gym/福利
Founding Software Engineer [33397]
Founding Software Engineer [33397]

Stealth Startup • San Francisco (CA)

On-site
USD 200,000 - 280,000
Research Engineer - Midtraining
Research Engineer - Midtraining

Periodic Labs • Menlo Park (CA)

On-site
USD 250,000 - 350,000
Research Engineer
Research Engineer

Cerebras • San Jose (CA)

On-site
USD 180,000 - 240,000
Founding Research Engineer, Model Training
Founding Research Engineer, Model Training

CellType Inc. • New York (NY)

On-site
USD 120,000 - 160,000
Founding Engineer - Post-training LLM's
Founding Engineer - Post-training LLM's

CT19 • New York (NY)

On-site
USD 180,000 - 260,000
Member of Technical Staff - Research
Member of Technical Staff - Research

Vals AI • San Francisco (CA)

On-site
USD 120,000 - 180,000
Relocation support
Health insurance
Lunch and snacks provided
+2
Member of Technical Staff, Post-Training & Applied Research
Member of Technical Staff, Post-Training & Applied Research

San Francisco Tensor Company • San Francisco (CA)

On-site
USD 275,000 - 315,000
Relocation assistance