Senior Data Engineer

SG Analytics

Mumbai Suburban

Hybrid

INR 1,200,000 - 1,800,000

Full time

13 days ago
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

SG Analytics, a Straive company, seeks a data engineering leader to design and govern metadata-driven data pipelines. You will build a Delta Lake-based lakehouse on Azure, ingest diverse data sources, and deliver governance, lineage, and secure access controls across the platform.

You will collaborate with AI developers and product owners to deliver scalable data products and RAG-enabled workflows, leveraging Azure AI services and best practices for data contracts.

Qualifications

  • Bachelor’s degree in Computer science, Engineering, Data Science, or a related technical field.
  • 5+ years hands-on experience building production data pipelines and data platforms.
  • Strong experience with Azure data stack (ADLS Gen2, Azure Databricks, Data Factory/Synapse).
  • Proficient in Python and SQL; Spark (PySpark) at scale; Delta Lake and lakehouse patterns.

Responsibilities

  • Design, build, and evolve a lakehouse architecture on Azure (ADLS Gen2, Databricks).
  • Develop robust ETL/ELT pipelines ingesting data from multiple sources.
  • Create metadata-driven ingestion frameworks and reusable connectors.
  • Model gold-layer data products for retrieval-augmented generation and RAG workflows.
  • Standardize on Delta Lake; explore Iceberg/UniForm for portability.
  • Enforce governance, lineage, access controls, and compliance (Purview, Unity Catalog).
  • Optimize performance and cost; ensure reliability, observability, and CI/CD.

Skills

Analytical thinking
Problem solving
Team collaboration
Communication

Education

Bachelor’s degree in Computer science, Engineering, Data Science, or related field

Tools

Azure Databricks
ADLS Gen2
Azure Data Factory
Synapse
Delta Lake
Parquet/Avro/JSON
dbt
Great Expectations
Terraform
Bicep
Azure DevOps
GitHub Actions
Purview
Unity Catalog

Job description

ABOUT US

SG Analytics (SGA), a Straive company, is a leading global data and AI consulting firm delivering solutions across AI, Data, Technology, and Research. With deep expertise in BFSI, Capital Markets, TMT (Technology, Media & Telecom), and other emerging industries, SGA empowers clients with Ins(AI)ghts for Business Success through data-driven transformation. A Great Place to Work certified company, SGA has a team of over 1,600 professionals across the U.S.A, U.K, Switzerland, Poland, and India. Recognized by Gartner, Everest Group, ISG, and featured in the Deloitte Technology Fast 50 India 2024 and Financial Times & Statista APAC 2025 High Growth Companies, SGA delivers lasting impact at the intersection of data and innovation.


JOB DESCRIPTION

This role is deliberately framed around orchestration, not assistance. You will not simply move data, you will design the metadata-driven, self-describing, governed architecture that lets AI systems reliably do the work. Your pipelines are the difference between an AI that guesses and an AI our professionals can stake their judgment on.

At SGA we look for individuals who welcome new ideas, encourage innovation, and are eager to make an impact. Whether youre starting out in your career or taking your next step as a seasoned professional, the SGA experience is one-of-a-kind. You can design a career youll love from top to bottom - we give you the tools you need to succeed and the autonomy to reach your goals.


ROLES & RESPONSIBILITIES
  • Modern lakehouse architecture. Design, build, and evolve the companys lakehouse on Azure Data Lake Storage (ADLS Gen2) and Azure Databricks, using a medallion architecture (bronze silver gold) with Delta Lake as the canonical open table format for ACID transactions, schema enforcement, and time travel.
  • ETL/ELT pipeline engineering. Build robust batch and streaming pipelines that ingest, cleanse, conform, and integrate data from firm source systems, client-provided data, and third-party feeds - favoring ELT-in-lakehouse patterns with Azure Data Factory, Databricks Workflows, and Structured Streaming.
  • Reusable, metadata-driven ingestion. Develop configuration-driven ingestion frameworks and reusable connectors so new sources onboard in days, not weeks - aligned to the firm’s reusable-capability taxonomy (primitive and business-function layers).
  • AI-ready data products. Curate and model gold-layer data products and vector-ready content that ground Retrieval-Augmented Generation and agentic workflows through Azure AI Search, ensuring chunking, embeddings, and freshness meet retrieval-quality standards.
  • Open table & file-format strategy. Standardize on Delta Lake while maintaining fluency across Parquet, Avro, ORC, and JSON at system boundaries; evaluate interoperability options (e.g., Apache Iceberg, UniForm) to keep the platform portable and future-proof.
  • Governance, lineage & compliance. Enforce end-to-end data governance, classification,

lineage, and access control through Microsoft Purview and Databricks Unity Catalog - with

controls that satisfy Section 7216, PCAOB, and the NIST AI Risk Management Framework,

including PII/PHI protection and client-data segregation.

  • Performance & cost optimization. Tune the platform for scale and spend through partitioning, liquid clustering / Z-ordering, file compaction, incremental and change-data-capture loads, and right-sized compute - treating cost per workload as a first-class engineering metric.
  • Reliability & observability. Establish CI/CD, automated data quality and contract testing, and pipeline observability (freshness, volume, schema drift, lineage) so data issues are detected before they reach an AI system or a client deliverable.
  • Cross-functional partnership. Collaborate closely with AI Developers, Product Owners,

Governance & Risk leaders, and Cloud Architects - and contribute to The Guild, technical

steering committee - to consolidate point solutions onto shared, governed platform

capabilities.


Basic Qualifications
  • Bachelor’s degree in Computer science, Engineering, Data Science, or a related technical field.
  • 5+ years relevant experience building production data pipelines and data platforms, with

hands-on ownership of ETL/ELT design and delivery.

  • Demonstrated experience on the Azure data stack, including ADLS Gen2 and Azure Databricks (Spark), plus Azure Data Factory and/or Synapse.
  • Strong programming in Python and SQL, with proven Spark (PySpark/Spark SQL) proficiency at scale.
  • Working expertise with Delta Lake and the medallion / lakehouse pattern, and fluency across columnar and row-based file formats (Parquet, Avro, JSON).
  • Practical experience implementing data governance, lineage, and access controls (e.g., Unity Catalog, Microsoft Purview).

Preferred/Desired Qualifications
  • Consulting or professional-services background with a strong bias for action.
  • Experience grounding AI/LLM systems with governed data - vector search, embeddings,

and RAG pipelines (Azure AI Search, Cosmos DB).

  • Streaming and CDC experience (Structured Streaming, Azure Event Hubs, Kafka, Debezium).
  • Transformation and data-contract tooling (dbt, Great Expectations, Delta Live Tables) and Infrastructure-as-Code (Terraform, Bicep).
  • Familiarity with open table-format interoperability (Apache Iceberg, Delta UniForm) and

lakehouse federation.

  • DevOps for data: Azure DevOps or GitHub Actions, Key Vault, Managed Identity, and cost

management tooling.

  • Exposure to regulated-data environments and standards relevant to tax, audit, and advisory (Section 7216, PCAOB, NIST AI RMF, SOC 2).

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Software Engineer, Data Engineering
Senior Software Engineer, Data Engineering

ValGenesis • Chennai

On-site
INR 1,000,000 - 1,500,000
Data Engineering Lead
Data Engineering Lead

Kumaran Systems • Hyderabad

On-site
INR 2,800,000 - 4,000,000
Technical Lead - Data Engineer ( Databricks Azure)
Technical Lead - Data Engineer ( Databricks Azure)

Srijan: Now Material • Gurugram District

On-site
INR 2,400,000 - 4,800,000
Data Engineer
Data Engineer

Coretek Services India • Hyderabad

Hybrid
INR 2,800,000 - 4,000,000
Azure Data Engineer
Azure Data Engineer

Artech L.L.C. • Dadri

On-site
INR 1,200,000 - 2,400,000
Lead Data Engineer
Lead Data Engineer

Keka Inc. • Dadri

On-site
INR 1,500,000 - 2,100,000
Senior Analyst - Data Engineering
Senior Analyst - Data Engineering

Darwinbox Digital Solutions Pvt. Ltd. • Bengaluru

On-site
INR 2,500,000 - 3,800,000
Azure Data Engineer/azure data architect/Lead data engineer
Azure Data Engineer/azure data architect/Lead data engineer

Tata Consultancy Services • Hyderabad, Chennai District

On-site
INR 1,500,000 - 2,100,000
Senior Data Engineer (India)
Senior Data Engineer (India)

Alimentiv • India

On-site
INR 3,000,000 - 5,000,000
AI Data Engineer
AI Data Engineer

SG Analytics • Mumbai

Hybrid
INR 1,500,000 - 2,700,000