Turn this role into an interview — a resume and cover letter built around what this employer wants.
Amgen seeks a Senior Data Scientist - Protein Data Pipelines to build scalable data pipelines for protein sequence, structure, and function modeling. You will enable ML-ready data assets, support model training, inference, deployment, and monitoring across research programs.
The role bridges data engineering, MLOps, computational biology, and applied ML, collaborating with ML developers and wet-lab teams to create robust, reusable data systems for discovery pipelines.
Research Job Description Senior Data Scientist - Protein Data Pipelines
Role Summary
The Senior Data Scientist - Protein Data Pipelines will play a critical role in enabling predictive modeling for protein sequence, structure, and function by building scalable, reliable, and reproducible data pipelines. This role will focus on transforming protein property data and related scientific outputs into ML-amenable assets that support model training, inference, deployment, and ongoing use across research programs. Working at the intersection of data engineering, MLOps, computational biology, and applied machine learning, this individual will partner with ML developers, wet-lab scientists, domain experts, and distributed technical teams to translate scientific and engineering needs into robust data and inference solutions. The successful candidate will develop reusable frameworks for data engineering, model inference, deployment, validation, testing, and monitoring across in-house and external machine learning models. This role is ideal for someone who enjoys building production-ready scientific data systems, collaborating across disciplines, and converting complex domain needs into maintainable technical solutions that scale across discovery pipelines.
Design and maintain scalable data pipelines that support predictive model training, with emphasis on protein sequence or structure-to-function applications.
Build ML-model amenable data assets for protein property data that are readable, quality-controlled, reproducible, and suitable for reuse across programs.
Translate scientific and engineering needs into reliable data solutions that support ongoing research and model-development workflows.
Develop deployment strategies and pipelines to embed trained models into ongoing projects.
Develop reusable inference, deployment, and testing frameworks for in-house and external machine learning models.
Convert model-development outputs into maintainable technical solutions that can be used reliably by research teams.
Establish data quality, validation, monitoring, and reproducibility practices for protein property and related scientific datasets.
Implement validation and monitoring approaches that improve confidence in downstream model training, inference, and deployment.
Document data lineage, assumptions, validation outcomes, and reproducibility practices to support long-term reuse.
Serve as a liaison between machine-learning developers and domain experts, including wet-lab collaborators where applicable.
Own and mediate collaborations between ML developers and wet-lab teams to ensure that data, modeling, and experimental needs are aligned.
Coordinate technical work across distributed teams and help align implementation plans, dependencies, and delivery timelines.
Document systems, pipeline behavior, operational expectations, and technical decisions to support adoption and maintenance.
Support knowledge sharing across research, ML, and engineering teams through clear documentation, examples, and reusable implementation patterns.
Scale data and modeling infrastructure practices across research programs and pipelines.
The ideal candidate combines strong data-engineering and MLOps expertise with enough scientific domain fluency to work effectively with ML developers, and experimental collaborators. They enjoy building reusable systems that make complex scientific data reliable, reproducible, and actionable for predictive modeling. Candidates may come from data science, data engineering, machine learning infrastructure, computational biology, computational chemistry, computational materials science, bioinformatics, or research informatics backgrounds. They are motivated by bridging scientific and engineering needs, supporting production-ready model use, and scaling technical solutions across discovery programs.
This role will build ML-amenable data pipelines for protein property data, mediate collaborations between ML developers and wet-lab teams, and scale data and modeling infrastructure across research programs and pipelines. By converting complex domain needs into maintainable, production-ready technical solutions, the Senior Data Scientist - Protein Data Pipelines will help accelerate reliable model development, deployment, and adoption across Large Molecule Discovery.
Amgen is committed to unlocking the potential of biology for patients suffering from serious illnesses by discovering, developing, manufacturing and delivering innovative human therapeutics. This approach begins by using tools like advanced human genetics to unravel the complexities of disease and understand the fundamentals of human biology.
Amgen focuses on areas of high unmet medical need and leverages its biologics manufacturing expertise to strive for solutions that improve health outcomes and dramatically improve people's lives.
A biotechnology pioneer since 1980, Amgen has grown to be one of the world's leading independent biotechnology companies, has reached millions of patients around the world and is developing a pipeline of medicines with breakaway potential.
For more information, visit www.amgen.com and follow us on www.twitter.com/amgen