A health technology firm in Chennai is seeking a mid-senior level Data Engineer to develop ETL pipelines and maintain data integrity using GCP Dataform and BigQuery. The role demands strong skills in R Shiny for UI development, experience with on-premise databases, and the ability to collaborate effectively within teams. This full-time position requires participation across various software development lifecycle phases and offers an opportunity to contribute to innovative health solutions.
Qualifications
Strong proficiency in R Shiny for UI development.
Experience designing, developing, and maintaining ETL pipelines using GCP.
Proven experience working with data manipulation and transformation in R.
Responsibilities
Develop and maintain ETL pipelines using GCP Dataform and BigQuery.
Implement data manipulation logic in R.
Collaborate with teams to ensure quality software delivery.
Participate in software development lifecycle phases.
Skills
R Shiny
Data Science
ETL Pipelines
Data Manipulation
Oracle Database
Git
Splunk
Google Cloud Platform (GCP)
Analytical Skills
Collaboration
Tools
GCP Dataform
BigQuery
Tableau
Job description
Responsibilities
Develop and maintain efficient and scalable ETL pipelines using GCP Dataform and BigQuery to extract, transform, and load data from various on‑premise (Oracle) and cloud‑based sources. This includes leveraging Big R query for accessing on‑premise Oracle data.
Develop and implement data manipulation and transformation logic in R, creating a longitudinal data format with unique member identifiers.
Develop and implement comprehensive logging and monitoring using Splunk.
Collaborate with other developers, data scientists, and stakeholders to ensure the timely delivery of high‑quality software.
Participate in all phases of the software development lifecycle, from requirements gathering and design to testing and deployment.
Contribute to the maintenance and improvement of existing application functionality.
Work within a Git‑based version control system.
Manage data in a dedicated GCP project, adhering to best practices for cloud security and scalability.
Contribute to the creation of summary statistics for groups via the Population Assessment Tool (PAT).
Required Skillsets
Strong proficiency in R Shiny for UI development.
Strong proficiency in data science background and Data Analyst background.
Proven experience in designing, developing, and maintaining ETL pipelines, preferably using GCP Dataform and BigQuery.
Experience with data manipulation and transformation in R, including creating longitudinal datasets.
Experience working with on‑premise databases (Oracle), preferably using Big R query for data access.
Experience with Git for version control.
Experience with Splunk for logging and monitoring.
Experience working with cloud platforms, specifically Google Cloud Platform (GCP).
Strong analytical and problem‑solving skills.
Excellent communication and collaboration skills.
Good to Have Skillsets
Experience with Tableau for dashboarding and data visualization.
Experience with advanced data visualization techniques.
Experience working in an Agile development environment.