An application made for this job — a tailored resume and cover letter that speak straight to the posting.
DatologyAI in San Mateo, CA is offering a Summer 2027 Research Intern role focused on investigating how intervention on training data can shape model behavior. You will work with a small team to explore practical improvements drawn from literature and experiments.
You will transform literature into actionable ideas, conduct high-risk high-reward research, and collaborate with engineers and customers to inform product direction.
Models are what they eat. But a large portion of training compute is wasted training on data that are already learned, irrelevant, or even harmful, leading to worse models that cost more to train and deploy.
Models are what they eat. But a large portion of training compute is wasted training on data that are already learned, irrelevant, or even harmful, leading to worse models that cost more to train and deploy.
At DatologyAI, we've built a state of the art data curation suite to automatically curate and optimize petabytes of data to create the best possible training data for your models. Training on curated data can dramatically reduce training time and cost (7-40x faster training depending on the use case), dramatically increase model performance as if you had trained on >10x more raw data without increasing the cost of training, and allow smaller models with fewer than half the parameters to outperform larger models despite using far less compute at inference time, substantially reducing the cost of deployment. For more details, check out our recent research on synthetic data scaling (BeyondWeb) and pretraining with domain-specific data (The Finetuner's Fallacy).
We raised a total of $57.5M in two rounds, a Seed and Series A. Our investors include Felicis Ventures, Radical Ventures, Amplify Partners, Microsoft, Amazon, and AI visionaries like Geoff Hinton, Yann LeCun, Jeff Dean, and many others who deeply understand the importance and difficulty of identifying and optimizing the best possible training data for models. Our team has pioneered this frontier research area and has the deep expertise on both data research and data engineering necessary to solve this incredibly challenging problem and make data curation easy for anyone who wants to train their own model on their own data.
This role is based in San Mateo, CA. We are in office 4 days a week.
The internship is for Summer 2027 and will take place sometime between May and August 2027.
As a Research Intern at DatologyAI, you will conduct research investigating how intervention on training data can improve the quality and shape the behavior of deep learning models.
Here is what your day-to-day would look like:
Ideal candidates should have strong coding skills with experience with one of the following: