Research Engineer

datologyai

San Mateo

On-site

PHP 11,314,000 - 18,856,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

100% health benefits
401(k) plan with 4% match
Unlimited PTO
Paid Parental Leave
Wellness stipend
Learning & development stipend
Daily lunches in office
Relocation assistance

Job summary

DatologyAI is hiring a Research Engineer to advance cutting-edge data curation and translate research into our core product pipeline. You will collaborate with researchers and engineers to scale pipelines for massive datasets and ensure production-ready systems.

The role emphasizes Python, PyTorch, and distributed compute, with opportunities to work on GPU clusters in a cloud environment. Competitive compensation and equity are offered.

Qualifications

  • 3+ years building ML systems, data infrastructure, or large-scale distributed applications.
  • Fluency in Python and hands-on experience with PyTorch.
  • Experience with distributed data processing and/or distributed training (Spark, Ray, Dask, Snowflake).
  • Comfort operating large-scale compute, including GPU clusters and cloud infrastructure.

Responsibilities

  • Build and scale data processing and curation pipelines over large language, vision, and multimodal datasets.
  • Design experimentation infrastructure to enable rapid, scalable research iterations.
  • Harden research results into production-grade data systems used by customers.
  • Profile and optimize large-scale training and data workloads for performance and cost.

Skills

Python fluency
PyTorch experience
Distributed systems
Software engineering fundamentals

Tools

PyTorch
Spark
Ray
Dask
Snowflake

Job description

About the Company

Models are what they eat. But a large portion of training compute is wasted training on data that are already learned, irrelevant, or even harmful, leading to worse models that cost more to train and deploy.

At DatologyAI, we’ve built a state of the art data curation suite to automatically curate and optimize petabytes of data to create the best possible training data for your models. Training on curated data can dramatically reduce training time and cost (7-40x faster training depending on the use case), dramatically increase model performance as if you had trained on >10x more raw data without increasing the cost of training, and allow smaller models with fewer than half the parameters to outperform larger models despite using far less compute at inference time, substantially reducing the cost of deployment. For more details, check out our recent research on synthetic data scaling (BeyondWeb) and pretraining with domain-specific data (The Finetuner’s Fallacy).

We raised a total of $57.5M in two rounds, a Seed and Series A. Our investors include Felicis Ventures, Radical Ventures, Amplify Partners, Microsoft, Amazon, and AI visionaries like Geoff Hinton, Yann LeCun, Jeff Dean, and many others who deeply understand the importance and difficulty of identifying and optimizing the best possible training data for models. Our team has pioneered this frontier research area and has the deep expertise on both data research and data engineering necessary to solve this incredibly challenging problem and make data curation easy for anyone who wants to train their own model on their own data.

This role is based in San Mateo, CA. We are in office 4 days a week.

About the Role

As a Research Engineer, you will play a crucial role in conducting and enabling cutting-\u2011edge research and translating it into our core product pipeline. You will work closely with other members of the technical staff to develop and improve state-of-the-\u2011art data curation strategies. Your technical skills will accelerate our research and ensure that our product remains at the forefront of innovation.

What You'll Work On
  • You'll build and scale the data processing and curation pipelines that operate over massive language, vision, and multimodal datasets, and make them fast, reliable, and cheap to run.

  • You'll design the experimentation infrastructure that lets scientists iterate quickly at scale, turning a good idea into a running experiment in hours.

  • Our work is guided by concrete customer needs and product outcomes. You'll take research results and harden them into production-\u201grade systems that customers depend on.

  • You'll profile and optimize large-scale training and data workloads, because at frontier scale, performance and cost are research constraints.

About You
  • 3+ years building ML systems, data infrastructure, or large-\u2016scale distributed applications.

  • Strong software engineering fundamentals, fluency in Python, and hands-\u2011on experience with PyTorch.

  • Experience with distributed data processing and/or distributed training, using tools such as Spark, Ray, Dask, or Snowflake.

  • Comfort operating large-\u2016scale compute, including GPU clusters and cloud infrastructure.

  • Enough machine learning depth to collaborate substantively with researchers, not just implement their specs.

  • A demonstrated track record of shipping systems that others rely on, whether through production infrastructure, open-\u2011source tools, or other artifacts.

Compensation

At DatologyAI, we are dedicated to rewarding talent with competitive salary and meaningful equity. The salary for this position ranges from $180,000 to $300,000.

  • Starting pay is based on job-\u2011related skills, experience, qualifications, and interview performance.

Benefits:

  • 100% covered health benefits (medical, vision, and dental).

  • 401(k) plan with a generous 4% company match.

  • Unlimited PTO policy

  • Paid Parental Leave of 12 weeks, plus 6 months of WFH flexibility.

  • Annual $2,000 wellness stipend.

  • Annual $1,000 learning and development stipend.

  • Daily lunches and snacks are provided in our office!

  • Relocation assistance for employees moving to the Bay Area.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Research Scientist
Research Scientist

datologyai • San Mateo

On-site
PHP 11,264,000 - 18,773,000
Health benefits
401(k) plan with company match
Unlimited PTO
+5
Software Engineer, Front-end
Software Engineer, Front-end

datologyai • San Mateo

Hybrid
PHP 11,285,000 - 18,809,000
Health benefits
401(k) plan with 4% match
Unlimited PTO
+5
Software Engineer, Infrastructure
Software Engineer, Infrastructure

datologyai • San Mateo

On-site
PHP 11,285,000 - 18,809,000
Health benefits (medical, vision, and牙
401(k) with 4% match
Unlimited PTO
+5
Software Engineer, Data Infrastructure
Software Engineer, Data Infrastructure

datologyai • San Mateo

On-site
PHP 11,285,000 - 18,809,000
Health benefits
401(k) with match
Unlimited PTO
+4
Forward Deployed AI Engineer (Post-Sales)
Forward Deployed AI Engineer (Post-Sales)

datologyai • San Mateo

On-site
PHP 14,375,000 - 18,750,000
100% covered health benefits
401(k) plan with 4% company match
Unlimited PTO policy
+4
Product Designer
Product Designer

datologyai • San Mateo

On-site
PHP 11,229,000 - 15,596,000
Health benefits (100% covered)
401(k) plan with 4% company match
Unlimited PTO
+4
Product Manager
Product Manager

datologyai • San Mateo

On-site
PHP 13,454,000 - 18,773,000
Health benefits
401(k) match
Unlimited PTO
+5
Software Engineer, Cloud Infrastructure
Software Engineer, Cloud Infrastructure

datologyai • San Mateo

On-site
PHP 11,285,000 - 18,809,000
Health benefits
401(k) plan with company match
Unlimited PTO
+5
AI Developer Experience & Media Lead
AI Developer Experience & Media Lead

datologyai • San Mateo

On-site
PHP 10,006,000 - 14,384,000
Health benefits
401(k) plan with company match
Unlimited PTO
+5
Software Engineer Intern, Infrastructure (Winter 2027)
Software Engineer Intern, Infrastructure (Winter 2027)

datologyai • San Mateo

On-site
PHP 2,093,000 - 3,139,000
Health benefits
401(k) plan with company match
Unlimited PTO
+5