Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.
Invok Hr in Pune is seeking a Principal Data Engineer to lead the Finance Data Hub and define six foundational engineering frameworks for data pipelines and platforms. The role emphasizes platform engineering, governance, and adoption of GenAI tools to accelerate data product delivery.
The ideal candidate has deep Snowflake expertise, strong SQL/Python skills, and experience building scalable data pipelines, data quality, and observability frameworks in a DataOps environment.
If you desire to be part of something special, to be part of a winning team, to be part of a fun team winning is fun. We are looking forward to hire Principal Data Engineer in Pune,
India. This exciting role offers opportunity to:
Requirement :
transformation, orchestration, and serving layers. Strong SQL and Python
proficiency
mentoring engineers, conducting design reviews, and influencing engineering
direction without necessarily holding a direct management title
tables, streams and tasks, RBAC design, row access policies, dynamic masking,
warehouse sizing, and query optimization. Snowflake certification is a strong plus
Actions for CI/CD automation. Experience designing branching strategies and
automated test/deploy pipelines for data workloads
packages, sources, and exposures. Coalesce experience or familiarity is an
advantage. Understanding of DAG-based transformation orchestration
Snowflake row access policies). Understands the intersection of data governance
policy and platform enforcement
instrumentation in production environments
(star schema, SCD types), and modern lakehouse/warehouse modeling patterns.
Has published or enforced modeling standards
re-platforming, cycle time reduction, or adoption of modern tooling. Can articulate
before/after outcomes with metrics
workflows AI code assistants, LLM-powered documentation, natural language
querying, or AI-driven anomaly analysis.
or transformation models. Understands test pyramid concepts in a data context:
unit, integration, and contract tests
Design, document, and version-control all six engineering frameworks in a central
standards repository (GitHub), ensuring they are discoverable, living documents with clear
change governance.
Conduct framework enablement sessions, workshops, and pair-programming to drive
active adoption not just publication across the engineering team.
Define conformance criteria and lightweight review checkpoints so that new pipeline work
is assessed against framework standards before promotion to production.
Act as the technical authority and tiebreaker on engineering design decisions
establishing consistent patterns while preserving pragmatic flexibility where needed.
Design and implement CI/CD pipelines for data engineering workloads using GitHub
Actions or equivalent - covering lint, unit test, schema validation, and environment
promotion stages.
Establish automated unit testing patterns - including test coverage standards and
coverage reporting.
Implement data contract frameworks at ingestion, transformation, and consumption
boundaries - defining schemas, SLOs, and acceptable value ranges as code.
Build reusable data quality monitoring templates - parameterizable and composable
across data products.
Instrument pipelines with observability metadata: lineage, runtime metrics, freshness
timestamps, and row count deltas - surfaced into operational dashboards.
Design and test the incident response workflow for data quality breaches: automated
alerting, quarantine patterns, stakeholder notification, and self-healing logic where
feasible.
Design and implement scalable RBAC models in Snowflake - covering functional roles,
object ownership hierarchies, and data product consumer roles.
Build row-level security (RLS) frameworks using Snowflake row access policies - creating
reusable, metadata-driven policy templates that can be applied consistently across
Finance data products.
Define and implement dynamic data masking policies aligned to the data classification
taxonomy - ensuring sensitive financial data is protected at the platform layer, not just the
application layer.
Govern Snowflake resource utilization: warehouse sizing standards, query optimization
guidelines, and cost attribution tagging by domain or product.
Champion the exploration and adoption of GenAI tooling to amplify data engineering
productivity - including AI-assisted SQL/python code generation, automated
documentation, and intelligent pipeline debugging.
Prototype and evaluate LLM-powered data engineering assistants: natural language to SQL
interfaces, automated data contract generation, and AI-driven anomaly root cause
analysis.
Define guardrails and governance standards for GenAI use in data engineering workflows -
covering code review requirements, hallucination risk in data contexts, and audit
traceability.
Share findings and tooling recommendations with the wider data engineering community
through internal demos, documentation, and engineering blog posts.
Identify and eliminate sources of engineering friction - legacy patterns, manual
deployment steps, inconsistent environments - and replace with automated, standardsdriven equivalents.
Measure and report on delivery cycle time improvements attributable to framework
adoption: pipeline build time, time to production, defect escape rate, and time to recovery.
Lead or contribute to data engineering modernization initiatives: migrating legacy ETL
workloads, re-platforming to Snowflake, and adopting modern orchestration patterns.