Role: Data Engineer
Experience Level: 2 to 4 Years
Work location: Airoli, Mumbai
What youll do:
You will be working as a Data Engineer within the healthcare domain, designing and delivering big data pipelines for structured and unstructured data that are running across multiple geographies, helping healthcare organizations achieve their business goals with use of data ingestion technologies, cloud services & DevOps. You will be working with Architects from other specialties such as Cloud engineering, Software engineering, ML engineering to create platforms, solutions and applications that cater to latest trends in the healthcare industry such as data intelligence platforms, enabling advanced analytics across pharmaceutical research, drug commercialization, sales & marketing, supply chain, and other key business domains, amongst others.
Role & Responsibilities:
- Work with cloud engineers and customers to solve complex big data problems by developing utilities for migration, storage, and processing on Azure Cloud.
- Design and build a cloud migration strategy for cloud and on-premise applications, leveraging Azure-native services.
- Develop, schedule, and manage data pipelines using Azure Data Factory (ADF) and Databricks Workflows.
- Design and implement data models following best practices including Star Schema, Snowflake Schema, and Slowly Changing Dimensions (SCDs – Type 1, 2, 3).
- Build and maintain data warehousing solutions on Azure Synapse Analytics / Databricks SQL Warehouse, ensuring data quality, consistency, and performance.
- Develop scalable ETL/ELT pipelines using PySpark on Databricks, processing terabytes to petabytes of data daily.
- Implement and manage Unity Catalog for fine-grained data governance, access control, lineage tracking, and data discovery across Databricks workspaces.
- Leverage Azure Data Lake Storage Gen 2 (ADLS Gen 2) for scalable, secure, and hierarchical data storage.
- Automate workflows and integrate systems using Azure Logic Apps for event-driven orchestration and business process automation.
- Diagnose and troubleshoot complex distributed systems problems and develop solutions with significant impact at massive scale.
- Communicate with a wide set of teams including Infrastructure, Network, Engineering, DevOps, and cloud customers.
- Build advanced tooling for automation, testing, monitoring, administration, and data operations across multiple cloud clusters.
Skills expectation:
- Must have:
- 3+ years of hands-on experience in data engineering with strong proficiency in PySpark and Python.
- Deep expertise in Databricks, including:
- Unity Catalog – data governance, access control, data lineage, and cataloging.
- Delta Lake – ACID transactions, time travel, and optimized storage.
- Databricks Workflows – pipeline orchestration and scheduling.
- Strong SQL and Advanced SQL skills for data transformation, aggregation, and performance tuning.
- Solid understanding of Data Modelling concepts:
- Dimensional modelling (Star & Snowflake schemas)
- Slowly Changing Dimensions (SCDs) – Type 1, 2, and 3
- Data Vault and Medallion Architecture (Bronze/Silver/Gold layers)
- Hands‑on experience with Data Warehousing principles and implementation.
- Proficiency in Azure Services:
- Azure Data Factory (ADF) – pipeline creation, triggers, linked services, and monitoring.
- Azure Data Lake Storage Gen 2 (ADLS Gen 2) – hierarchical namespace, access control, and integration.
- Azure Logic Apps – workflow automation and system integration.
- Experience building and supporting large-scale data systems in production environments.
- Familiarity with NoSQL databases and cloud-native data services.
- Strong understanding of distributed systems and data structures.
- Good to have:
- Experience working with healthcare and pharmaceutical Data
- Designing and development of ETL pipeline
- Requirement gathering and understanding of the problem statement
- End-to-end ownership of the entire delivery of the project
- Designing and documentation of the solution
- Team management and mentoring
Leadership qualities
- Ability to lead technology teams and provide them mentorship / support to accelerate performance.
- Experience in leading multiple large projects as well as a deep understanding of Agile developments
- Effective communication with all the stakeholders involved.
- Communicate clearly about complex subjects and technical plans with technical and non technical audiences.
Behavioural skills:
- Effective communication with all the stakeholders involved
- Need to communicate clearly about complex subjects and technical plans
- Ability to mentor and groom junior members of the team and provide them with guidance and roadmap.
- Must also have the ability to interact with other members of the team (Juniors, ATA, TA, BA, etc) to get their designs from concept to development
Keeping various audiences in mind, engineers must write their reports in clear language accessible to all.
What is in it for you:
- Opportunity to learn cloud native services and how to utilize those services to solve various business problems.
- More hands‑on learning opportunity in Databricks, Python, SQL, etc
- Sponsored certification opportunity for various courses of your choice (eg, GCP, AWS, Azure, Tableau, Looker,etc)
- Allocated work will allow you to not only just complete the task but you will be given an opportunity to own-up the work and take responsibility of it’s end-to-end delivery