Senior Data Scientist - Protein Data Pipelines

Biopharma Careers

Hyderabad

On-site

INR 3,500,000 - 5,200,000

Full time

4 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Biopharma Careers in Hyderabad seeks a Senior Data Scientist - Protein Data Pipelines to build scalable data pipelines for protein properties and ML-ready assets. You will transform protein data into reliable, reproducible pipelines supporting model training, deployment, and ongoing use across research programs.

The role emphasizes cross-functional collaboration with ML developers and wet-lab teams, data quality, reproducibility, and scalable infrastructure across discovery projects.

Qualifications

  • Bachelor's degree with relevant professional experience in computational biology, bioinformatics, or related field.
  • Master's degree with 4+ years of relevant experience, or PhD.
  • Experience building scalable data pipelines and ML workflows for biological data.

Responsibilities

  • Design and maintain scalable data pipelines to support predictive model training for protein sequence/structure-to-function tasks.
  • Build ML-model amenable data assets that are readable, quality-controlled, and reusable.
  • Translate scientific needs into reliable data solutions for research and model-development workflows.
  • Develop deployment strategies and pipelines to embed trained models into ongoing projects.
  • Create reusable inference, deployment, and testing frameworks for ML models.

Skills

Python
SQL
Data pipelines
MLOps
MLflow
Data quality
Documentation

Education

Bachelor's degree in Computational Biology, Bioinformatics, Life Sciences, Computational Chemistry, Chemical Engineering, Materials Science, Data Science, or related field
Master's degree
PhD

Tools

Databricks

Job description

Role Summary

The Senior Data Scientist - Protein Data Pipelines will play a critical role in enabling predictive modeling for protein sequence, structure, and function by building scalable, reliable, and reproducible data pipelines. This role will focus on transforming protein property data and related scientific outputs into ML-amenable assets that support model training, inference, deployment, and ongoing use across research programs.

Key Responsibilities
Scalable Data Pipelines for model training
  • Design and maintain scalable data pipelines that support predictive model training, with emphasis on protein sequence or structure-to-function applications.
  • Build ML-model amenable data assets for protein property data that are readable, quality-controlled, reproducible, and suitable for reuse across programs.
  • Translate scientific and engineering needs into reliable data solutions that support ongoing research and model-development workflows.
Model Deployment & Inference Pipelines
  • Develop deployment strategies and pipelines to embed trained models into ongoing projects.
  • Develop reusable inference, deployment, and testing frameworks for in-house and external machine learning models.
  • Convert model-development outputs into maintainable technical solutions that can be used reliably by research teams.
Data Quality, Validation & Reproducibility
  • Establish data quality, validation, monitoring, and reproducibility practices for protein property and related scientific datasets.
  • Implement validation and monitoring approaches that improve confidence in downstream model training, inference, and deployment.
  • Document data lineage, assumptions, validation outcomes, and reproducibility practices to support long-term reuse.
Cross-Functional Collaboration & Technical Coordination
  • Serve as a liaison between machine-learning developers and domain experts, including wet-lab collaborators where applicable.
  • Own and mediate collaborations between ML developers and wet-lab teams to ensure that data, modeling, and experimental needs are aligned.
  • Coordinate technical work across distributed teams and help align implementation plans, dependencies, and delivery timelines.
Documentation & Knowledge Sharing
  • Document systems, pipeline behavior, operational expectations, and technical decisions to support adoption and maintenance.
  • Support knowledge sharing across research, ML, and engineering teams through clear documentation, examples, and reusable implementation patterns.
  • Scale data and modeling infrastructure practices across research programs and pipelines.
Basic Qualifications

Bachelor’s degree in Computational Biology, Bioinformatics, Life Sciences, Computational Chemistry, Chemical Engineering, Materials Science, Data Science, or a related quantitative field and relevant professional experience.

Experience Requirements
  • Bachelor’s degree and 6+ years of relevant experience, OR
  • Master’s degree and 4+ years of relevant experience, OR
  • PhD
Preferred Qualifications
Scalable Data Engineering
  • Strong experience building scalable data pipelines in Python and/or SQL.
  • Experience designing readable, reusable, and maintainable data-processing workflows for scientific or machine-learning applications.
  • Experience with data pipeline automation, preferably using Databricks.
MLOps, Inference & Deployment
  • Hands-on experience owning reusable, end-to-end MLOps for at least one machine learning model.
  • Experience with MLflow is preferred; experience with other model-lifecycle, deployment, or tracking frameworks is also welcome.
  • Experience developing deployment, inference, validation, or testing workflows that support production-like use of machine learning models.
Data Quality, Monitoring & Reproducibility
  • Knowledge of data quality control, validation, and monitoring practices.
  • Experience applying reproducibility practices to scientific data, model-training datasets, or inference workflows.
  • Ability to identify data quality risks and develop practical controls for downstream model use.
Scientific Domain Experience
  • Familiarity with computational biology, computational chemistry, computational materials science, or related fields.
  • Experience working with protein sequence, protein structure, protein property, or related scientific datasets is beneficial.
  • Preferred experience collaborating with wet-lab teams and translating experimental needs into data or modeling workflows.
Communication & Collaboration
  • Ability to communicate effectively with machine-learning developers, software and data engineers, domain experts, and research scientists.
  • Experience coordinating technical work across distributed or cross-functional teams.
  • Strong documentation habits and commitment to knowledge sharing, maintainability, and long-term adoption.
Success Measures
  • Delivery of readable, quality-controlled, and reproducible data pipelines for protein property data.
  • Successful embedding of trained models into ongoing projects through reliable deployment and inference pipelines.
  • Increased reuse of data-engineering, inference, deployment, validation, and testing frameworks across research programs.
  • Improved confidence in model-training and inference data through practical quality, monitoring, and reproducibility practices.
  • Effective collaboration between ML developers, wet-lab teams, domain experts, and distributed technical partners.
  • Expansion of scalable data and modeling infrastructure across research programs and pipelines.
Typical Candidate Profile

The ideal candidate combines strong data-engineering and MLOps expertise with enough scientific domain fluency to work effectively with ML developers, and experimental collaborators. They enjoy building reusable systems that make complex scientific data reliable, reproducible, and actionable for predictive modeling.

Candidates may come from data science, data engineering, machine learning infrastructure, computational biology, computational chemistry, computational materials science, bioinformatics, or research informatics backgrounds. They are motivated by bridging scientific and engineering needs, supporting production-ready model use, and scaling technical solutions across discovery programs.

Organizational Impact

This role will build ML-amenable data pipelines for protein property data, mediate collaborations between ML developers and wet-lab teams, and scale data and modeling infrastructure across research programs and pipelines. By converting complex domain needs into maintainable, production-ready technical solutions, the Senior Data Scientist - Protein Data Pipelines will help accelerate reliable model development, deployment, and adoption across Large Molecule Discovery.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Data Scientist - Protein Data Pipelines
Senior Data Scientist - Protein Data Pipelines

Amgen Inc. (IR) • Hyderabad

On-site
INR 2,500,000 - 4,500,000
Senior Data Scientist - Protein Structure ML models
Senior Data Scientist - Protein Structure ML models

Biopharma Careers • Hyderabad

On-site
INR 3,000,000 - 4,200,000
Senior Data Scientist
Senior Data Scientist

Biopharma Careers • Hyderabad

On-site
INR 1,500,000 - 2,700,000
Senior Machine Learning Engineer
Senior Machine Learning Engineer

Biopharma Careers • Hyderabad

On-site
INR 2,000,000 - 3,200,000
Senior Machine Learning Engineer
Senior Machine Learning Engineer

Amgen Inc • Hyderabad

On-site
INR 2,800,000 - 5,500,000
Senior Machine Learning Engineer
Senior Machine Learning Engineer

Amgen • Hyderabad

On-site
INR 4,000,000 - 6,000,000
Data Engineer
Data Engineer

Excelra • Hyderabad, Bengaluru, Delhi

On-site
INR 1,200,000 - 1,800,000
Senior Systems Engineer - Data DevOps/MLOps
Senior Systems Engineer - Data DevOps/MLOps

EPAM Systems • Chennai District

On-site
INR 5,000,000 - 7,500,000
Senior Systems Engineer - Data DevOps/MLOps
Senior Systems Engineer - Data DevOps/MLOps

Epam Systems • Coimbatore District

On-site
INR 350,000 - 700,000
ML Ops Engineer - Generative AI, Digital Automation, & Integration
ML Ops Engineer - Generative AI, Digital Automation, & Integration

Biotale • India

Hybrid
INR 800,000 - 1,200,000