Join to apply for the Senior Data Analyst (DevOps Engineer) role at Novartis India
1 day ago Be among the first 25 applicants
Join to apply for the Senior Data Analyst (DevOps Engineer) role at Novartis India
Summary
Novartis Biomedical Research is seeking an experienced and highly motivated data and MLOps scientist/engineer to help us push the frontiers of data science and machine learning for Life Sciences and drug discovery. Within the Research Informatics division, you will take on a hands-on data scientist role at the intersection of science, R&D, and real-world impact. You will be part of a unique organization with an inter-disciplinary team at the forefront of AI/ML in drug discovery. You will have the opportunity to shape the next wave of drug discovery, working with diverse data modalities including omics datasets (genomics, transcriptomics, proteomics), spatial omics, compound structures, protein sequences, cellular measurements, safety studies, histopathology, clinical imaging, and readouts. This is a senior individual contributor role requiring demonstrated hands-on experience in Python, R, CLI/shell scripting, and diverse ML workflows, with the ability to perform under minimal supervision in a collaborative environment.
About The Role
Key Responsibilities:
- Collaborate closely with data scientists and subject-matter experts to fulfill data and computational needs.
- Validate and ensure the accuracy and quality of data by cleaning, shaping, analyzing, normalizing, and conforming it to existing models and vocabularies.
- Identify and rectify data inconsistencies and irregularities.
- Design data models and prepare data artifacts to effectively meet business needs.
- Promote a culture of transparency and communication regarding data modifications, lineage, and definitions to all stakeholders.
Essential Requirements
- Proficiency in Python and/or R, and other scripting languages.
- Experience with HPC, cloud workflows (AWS), setup, and deployments.
- Experience deploying ML models and resolving package dependencies from sources like GitHub, Hugging Face, etc.
- Experience with DevOps and MLOps practices.
- Data management expertise in relational, document, column, and graph datastores.
- Building ETL processes in high-performance environments like Databricks, AWS, Snowflake.
- Experience with ML frameworks such as PyTorch, TensorFlow is a plus.
- Knowledge of bioinformatics tools for sequence matching, alignment, clustering, data imputation, and visualization methods.
- Strong skills in extracting relevant data from diverse sources (Excel, PowerPoint, CSV, databases) with varying formats and potential missing or mislabeled data.
- Experience working with subject matter experts to harmonize datasets for machine learning.
- Experience in document mining and processing diverse data sources.
- Experience working with large public scientific datasets, preferably biology-related, is a plus.
- Ability to validate data accuracy and quality through cleaning, shaping, analyzing, and normalizing.
- Interest in current literature and applications of ML and data science in biological sciences.
- Understanding of common ML concepts such as training vs. test sets, overfitting, bias, feature extraction, LLMs, classifiers.
- Excellent English communication skills, proactive in seeking clarifications.
Desirable Requirements
- BSc in Computer Science, Informatics, or similar, or equivalent practical experience.
- Fluency in English.
Additional information about benefits, diversity and inclusion, accommodations, and how to connect with Novartis is provided in the original description.