Research Scientist - VLM Pretraining

Epsilon Labs, Inc.

San Francisco (CA)

On-site

USD 180,000 - 260,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Epsilon Labs, Inc. is seeking a Research Scientist to lead pretraining and scaling of vision-language models for radiology, across X-ray, CT, and MRI data. You will own the pretraining stage, design architectures, manage data mixtures, and drive model quality through evaluation.

You will work with PyTorch/JAX, large-scale distributed training, and contribute to publications and practical deployment in healthcare settings.

Qualifications

  • 6+ years of academia/industry experience in vision-language modeling or related fields.
  • Deep expertise in pretraining large vision-language models such as LLaVA, Flamingo, CogVLM, Qwen-VL, InternVL, or similar architectures.
  • Strong foundation in modern VLM pretraining techniques including connector architectures, high-resolution handling, and embedding spaces.

Responsibilities

  • Design, train, and scale vision-language foundation models for radiology applications, owning pretraining end-to-end.
  • Develop architectures suited to medical imaging with high-resolution tiling and token budgets.
  • Build and tune multimodal pretraining mixtures across captioning, VQA, grounding, and retrieval tasks.
  • Develop fine-grained visual grounding during pretraining with bounding boxes or masks.
  • Own pretraining evaluation (zero- and few-shot transfer, probing, and fine-tunability).
  • Train joint vision-language embedding spaces using contrastive and generative objectives.

Skills

Vision-language modeling
Multimodal learning
Pretraining large models
Medical imaging
PyTorch
JAX
Distributed training
Long-context training
Grounding
Production-quality code

Tools

PyTorch
JAX

Job description

About Us

We're tackling one of healthcare's most critical challenges in medical imaging and diagnostics. Our company operates at the intersection of cutting-edge AI and clinical practice, building technology that directly impacts patient outcomes. We've assembled one of the industry's most comprehensive and diverse medical imaging datasets and have a proven product-market fit with a substantial customer pipeline already in place.

Role Overview

We're seeking a Research Scientist with deep expertise in large-scale vision-language pretraining to join our ML Research team. You'll be at the forefront of developing state-of-the-art multimodal models for clinical use in radiology settings. This role owns the pretraining stage of our radiology report generation model: VLM architecture design, multimodal data and task mixtures, and the large-scale training runs that build grounded visual understanding across X-rays, CT scans, and MRI. You'll work with one of the largest and most diverse medical imaging datasets in the industry, paired with the reports that make multimodal pretraining at this scale possible, while maintaining the clinical rigor required for healthcare deployment. Post-training and RL are owned by a partner role you'll collaborate with closely.

Key Responsibilities
  • Design, train, and scale vision-language foundation models for radiology applications, owning the pretraining stage end to end.

  • Develop VLM architectures suited to medical imaging, including native and variable resolution handling, high-resolution tiling, connector design, and token budgets for volumetric studies.

  • Build and tune multimodal pretraining mixtures across captioning, VQA, grounding, and retrieval tasks, balancing data sources to avoid regressions in language capability.

  • Develop fine-grained visual grounding during pretraining, enabling models to localize findings within medical images using bounding boxes or segmentation masks.

  • Own pretraining evaluation (zero- and few-shot transfer, probing, and downstream fine-tunability) — as the signal for base model quality.

  • Train joint vision-language embedding spaces using contrastive and generative objectives, including region- and sentence-level alignment between images and reports.

  • Contribute hands‑on to all stages of pretraining including dataset curation, architecture design, distributed training, and handoff of base checkpoints to post‑training.

  • Stay current with cutting‑edge research in vision‑language modeling and large‑scale multimodal pretraining.

  • Drive research and technical excellence through conference publications and technical blog posts, establishing best practices for pretraining medical VLMs at scale.

Qualifications
  • 6+ years of academia/industry experience in vision-language modeling, multimodal learning, or related fields

  • Deep expertise in pretraining large vision-language models (e.g., LLaVA, Flamingo, CogVLM, Qwen-VL, InternVL, or similar architectures)

  • Strong foundation in modern VLM pretraining techniques including:

    • Vision-language connector and fusion architectures (projection, cross-attention, resampler-based)

    • Variable and high-resolution image handling (native resolution, dynamic tiling, token compression)

    • Contrastive and generative objectives for learning joint vision-language embedding spaces

    • Data and task mixture design, including curriculum and mixture-ratio ablations

  • Experience with fine-grained visual grounding (referring expression comprehension, phrase grounding, box or mask prediction)

  • Track record of implementing complex models from research papers and adapting them to new domains

  • Proficiency in PyTorch or JAX, with experience training large models on multi-GPU/distributed systems

  • Experience with autoregressive language modeling and long-context training

  • Hands‑on experience with medical imaging applications, particularly radiology report generation

  • Strong software engineering skills and ability to write production-quality code

Preferred Qualifications
  • Publications at top-tier conferences (NeurIPS, ICML, ICLR, CVPR, ACL, EMNLP, MICCAI)

  • Experience training vision encoders from scratch, or co‑designing them with a downstream VLM

  • Experience with interleaved image-text pretraining and synthetic recaptioning pipelines

  • Experience with 3D medical image processing and temporal modeling

  • Familiarity with clinical NLP and medical knowledge representation

  • Knowledge of evaluation methodologies for long‑form generation, including factuality assessment and hallucination detection

  • Experience with model interpretability, explainability, and uncertainty quantification in safety‑critical applications

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Research Scientist - VLM Pretraining
Research Scientist - VLM Pretraining

Epsilon Health • San Francisco (CA)

On-site
USD 180,000 - 250,000
Research Scientist - Post-training / RL
Research Scientist - Post-training / RL

Epsilon Labs, Inc. • San Francisco (CA)

On-site
USD 180,000 - 240,000
Lead VLM Pretraining Scientist for Medical Imaging
Lead VLM Pretraining Scientist for Medical Imaging

Epsilon Health • San Francisco (CA)

On-site
USD 180,000 - 250,000
Research Scientist - Vision Foundation Models
Research Scientist - Vision Foundation Models

Epsilon Labs, Inc. • San Francisco (CA)

On-site
USD 180,000 - 260,000
Research Scientist - Vision Foundation Models
Research Scientist - Vision Foundation Models

Epsilon Health • San Francisco (CA)

On-site
USD 180,000 - 240,000
Research Engineer - Data Quality & Evals
Research Engineer - Data Quality & Evals

Epsilon Health • San Francisco (CA)

On-site
USD 120,000 - 170,000
Senior Machine Learning Engineer
Senior Machine Learning Engineer

Clearview AI • United States

On-site
USD 180,000 - 250,000
Medical plans
Dental plans
Vision plans
+1
Radiology Vision-Language Pretraining Scientist
Radiology Vision-Language Pretraining Scientist

Epsilon Labs, Inc. • San Francisco (CA)

On-site
USD 180,000 - 260,000
Vision-Language Models (VLMs)
Vision-Language Models (VLMs)

TalentOla • Waukesha (WI)

On-site
USD 120,000 - 150,000
Senior ML Research Scientist
Senior ML Research Scientist

Rad AI • United States

On-site
USD 140,000 - 230,000