Lead Engineer, Data Engineering

Trinity Life Sciences

Bengaluru

On-site

INR 1,800,000 - 3,000,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Trinity Life Sciences Bengaluru is seeking a Senior Data Engineer to design and build scalable data pipelines using PySpark, Python and SQL on cloud platforms. You will translate architecture and design patterns into concrete, reusable workflows and ensure datasets are well-documented and trusted for analytics.

You will own Airflow DAGs, implement data quality checks, optimize jobs on Databricks and Snowflake, and collaborate with Analytics, Data Science, and Product teams to deliver reliable,

Qualifications

  • 6–9 years of production-grade data engineering experience.
  • Strong PySpark, SQL, and data modeling skills.
  • Experience with Airflow DAGs and cloud data warehouses.
  • Familiarity with life sciences commercial datasets.

Responsibilities

  • Build scalable batch and near real-time data pipelines.
  • Translate architecture and design patterns into reusable workflows.
  • Develop and maintain Airflow DAGs with monitoring for fault-tolerant pipelines.
  • Implement data quality checks and surface DQ metrics.
  • Collaborate with Analytics, Data Science, and Product teams.

Skills

PySpark
Airflow
SQL
Python
Data pipelines
Git
CI/CD

Tools

Databricks
Snowflake
BigQuery
Redshift
Delta Lake

Job description

  • Implement scalable batch and near real-time data pipelines using PySpark, Python, and SQL on cloud-native platforms.
  • Translate architecture and design patterns from the Data Engineering Principal into concrete, well-structured workflows and reusable components.
  • Develop and maintain Airflow DAGs, managing dependencies, schedules, and operational monitoring for robust, fault-tolerant pipelines.
  • Build data quality checks into pipelines (schema validation, record-level rules, thresholds, anomaly detection), and surface DQ metrics via dashboards or alerts.
  • Ingest, normalize, and harmonize life sciences commercial datasets (e.g., claims, prescription, sales, roster/territory, CRM, specialty pharmacy, payer/plan, formulary) into curated layers.
  • Implement data transformations and business rules to support use cases such as targeting, incentive compensation, call planning, patient journey, and market access analytics.
  • Work across data platforms such as Databricks and Snowflake, optimizing jobs through partitioning, clustering, caching, and cost/performance tuning.
  • Contribute to and follow data modeling standards (dimensional models, star/snowflake schemas) aligned to life sciences commercial subject areas.
  • Embed testing (unit tests for transformations, integration tests for pipelines, and regression checks on key metrics) into the development lifecycle.
  • Use version control and CI/CD workflows (e.g., Git, Jenkins, cloud-native tools) to promote reliable, repeatable deployments.
  • Collaborate with Analytics, Data Science, and Product teams to understand requirements and ensure datasets are usable, documented, and trusted
  • Produce clear technical documentation for pipelines, schemas, and business logic, making it easy for others to extend and maintain your work.
  • Participate in Agile ceremonies, provide realistic estimates, and own stories from design through to production support and handover.
What You Bring
  • 6-9 years of experience building production-grade data engineering solutions, with at least 3+ years hands-on with PySpark and Airflow.
  • Strong, hands-on PySpark skills for ETL/ELT, including working with large datasets, optimization (partitioning, joins, caching), and troubleshooting performance issues.
  • Solid Airflow experience: DAG design, scheduling, sensors, operators, and managing operational run health.
  • Practical experience with core life sciences commercial datasets: claims (medical/pharmacy), prescription (TRx/Nrx), sales, affiliations/rosters, call activity, and payer/plan/formulary data.
  • Strong SQL skills (complex joins, window functions, CTEs, performance tuning) across cloud data warehouses such as Snowflake, Redshift, BigQuery, or Databricks SQL.
  • Experience implementing data quality checks and validation patterns (row counts, referential integrity, business rules, reconciliation against source systems).
  • Familiarity with modern data platforms (Databricks, Snowflake, Delta Lake, S3/ADLS/GCS) and data lakehouse concepts.
  • Experience with at least one major cloud provider (AWS, Azure, or GCP) and core services used in data pipelines (e.g., S3/ADLS, Glue/Data Factory, Lambda/Functions).
  • Comfort working in Git-based workflows and using CI/CD for data pipelines.
  • Solid understanding of dimensional modeling and how to design fact and dimension tables for reporting and analytics.
  • Strong debugging skills, willingness to dive into logs and metrics, and bias toward shipping working solutions quickly and iterating.
  • Clear, concise communication and the ability to work with cross-functional stakeholders (consultants, analysts, data scientists) in a fast-paced environment.
Bonus Points
  • Deeper experience with pharma/life sciences commercial analytics (field force effectiveness, patient journey, HUB/SP, market access, or payer analytics).
  • Experience with dbt or similar tools for modular SQL transformations and documentation.
  • Exposure to data observability tools (e.g., Great Expectations, Soda, Monte Carlo) and data lineage/metadata tools.
  • Familiarity with Kafka or Kinesis for streaming use cases.
  • Experience containerizing data workloads (Docker) and running them on Kubernetes or similar orchestration platforms.
  • Prior consulting or client-facing work where you gathered requirements and translated them into concrete data solutions.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Data Engineer
Senior Data Engineer

Jobtailor • Hyderabad

On-site
INR 1,500,000 - 2,100,000
Senior Data Engineer
Senior Data Engineer

GAVS Technologies N.A., Inc • Chennai District

On-site
INR 1,000,000 - 2,000,000
Data Engineer – Databricks, Python, SQL & Cloud Migration
Data Engineer – Databricks, Python, SQL & Cloud Migration

Qloron Pvt Ltd • Hyderabad

On-site
INR 1,200,000 - 1,800,000
Senior Developer, Data Engineer
Senior Developer, Data Engineer

ICE • Hyderabad

On-site
INR 1,800,000 - 3,000,000
Senior Developer, Data Engineer
Senior Developer, Data Engineer

ICE Clear Europe Limited • Hyderabad

On-site
INR 1,500,000 - 2,000,000
Senior Azure Data Engineer
Senior Azure Data Engineer

Agilisium Technologies LLC. • Chennai

On-site
INR 800,000 - 1,200,000
JD14 - Consultant - DBT
JD14 - Consultant - DBT

Kivitronicsconsulting • Chennai District

On-site
INR 1,200,000 - 1,900,000
Data Analytics Engineer
Data Analytics Engineer

Monocept • Gurugram District

On-site
INR 1,200,000 - 1,800,000
Data Engineering Pipeline Engineer – Role Description
Data Engineering Pipeline Engineer – Role Description

Innoventes Technologies • Bengaluru

On-site
INR 1,000,000 - 1,500,000
Data Analytics Engineer
Data Analytics Engineer

EXL • Hyderabad, Pune District, Bengaluru

Hybrid
INR 1,200,000 - 1,800,000