- Assist in designing, developing, and testing scalable data pipelines for ingestion and transformation from multiple data sources into enterprise data lakes and warehouses
- Work with data engineers and scientists to clean, prepare, and analyze structured and unstructured data
- Support development and automation of model training, evaluation, and deployment pipelines
- Explore and analyze pharmaceutical commercial datasets for data-driven insights
- Participate in designing visualizations and dashboards for internal stakeholders
- Learn and apply hypothesis testing, regression, and classification
- Collaborate on implementing and monitoring ML models in production environments
- Document technical processes, models, and tools
- Experiment with data engineering and MLOps tools and techniques
- Participate in agile ceremonies, including sprint planning and retrospectives
Requirements
- Expected completion of one year of study from an accredited college or university before the internship begins
- Expected continued enrollment in an accredited college or university after the internship
- Student must be located in the United States for the duration of the internship
- Must be available to work a full-time schedule
- Must be available to accept and commit to future full-time employment by July 2028, if offered
- Must be 18 years or older
- Currently enrolled in a full-time Bachelor’s Degree program from an accredited college or university
- Minimum 3.0 GPA or equivalent
- Degree concentration in Information Technology, Computer Science, Engineering, Business, or a related field preferred
- Biotechnology, pharmaceutical, or health care industry experience preferred
- Intermediate knowledge of Microsoft Word, Excel, and PowerPoint preferred
- Foundational experience in Python or R, including pandas, NumPy, matplotlib, seaborn, scikit-learn, or XGBoost preferred
- Exposure to AWS, GCP, or Azure preferred
- Basic understanding of SQL and experience querying relational or big data sources preferred
- Interest or coursework in NLP, time-series analysis, or statistical modeling preferred
- Familiarity with Git and collaborative coding practices preferred
- Awareness of Databricks, Apache Spark, or Apache Airflow is a plus
- Coursework in business systems analysis, Lean, Agile/SCRUM, SDLC processes, software development, database modeling, web design and development, IT management, cloud-based application management, B2B collaboration, or enterprise systems preferred
- Must be authorized to work in the U.S. for the duration of the program
- Sponsorship for future full-time roles is not guaranteed
Core Competencies
Demonstrates foundational experience in Python or R for data analysis and model development, along with a basic understanding of SQL for querying data sources. Familiarity with data engineering and MLOps tools, as well as exposure to cloud platforms like AWS, GCP, or Azure, is essential for supporting scalable data pipelines and machine learning models.
Highest-signal resume keywords
- Python Programming
- Data Pipeline Development
- Machine Learning Model Deployment
- SQL Querying
- Data Visualization
ATS Optimization Keywords
Hard Skills
- Python
- R
- SQL
- Data Analysis
- Machine Learning
- Statistical Modeling
- Data Engineering
- Hypothesis Testing
- Regression
- Classification
Industry Keywords
- Pharmaceutical
- Biotechnology
- Health Care
- Agile
- Lean
- SDLC
- Business Systems Analysis
- Cloud-Based Application Management
- Enterprise Systems
- B2B Collaboration
Tools & Technologies
- Microsoft Word
- Microsoft Excel
- Microsoft PowerPoint
- AWS
- GCP
- Azure
- Git
- Databricks
- Apache Spark
- Apache Airflow