Senior Data Engineer

Confidential

Bengaluru

On-site

INR 1,800,000 - 2,400,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Hybrid options
Learning budget
Competitive pay
GPU infrastructure access

Job summary

Confidential, a mid-sized IT and engineering services company in Bengaluru, seeks a Senior Data Engineer to lead data solutions from ingestion to feature delivery. You will mentor engineers, drive strategy, and ensure production-grade pipelines with privacy and compliance baked in.

You will own stream and batch ingestion, build feature stores, and collaborate with data scientists to prevent training-serving skew while enforcing data contracts and SLAs.

Qualifications

  • Hands-on data engineering with production systems.
  • Experience building streaming pipelines at scale.
  • Familiarity with GDPR/CCPA and privacy by design.
  • Strong SQL and Python in production-grade code.
  • Experience with data lakehouse concepts (Delta Lake/ Iceberg).
  • Knowledge of ETL/ELT and data modeling basics.

Responsibilities

  • Design and implement governed ingestion pipelines.
  • Build streaming pipelines with schema validation and data quality checks.
  • Orchestrate batch and streaming ingestion with Airflow or equivalents.
  • Enforce privacy flags and data anonymisation at ingestion.
  • Maintain raw event store and deterministic identity logic.
  • Collaborate with ML engineers to align feature pipelines.
  • Define SLAs, monitor pipeline freshness, and alert on breaches.
  • Ensure data contracts and schema evolution across layers.

Skills

Data engineering
Kafka
Python
SQL
Airflow
Streaming pipelines
Data quality
Entity stitching

Tools

Kafka
Kinesis
Airflow
Databricks
Delta Lake
Iceberg
Great Expectations
dbt

Job description

  • Mid-Sized Pioneering IT and Engineering Services Company
  • Domains: Hi-Tech, Automotive, Manufacturing, Telecom, Medical and Life Sciences, Pharmaceutical
  • Successfully service Fortune 500 Companies
  • Customer Geographies: North America, Europe, Japan, Korea, China

Job Description – Senior Data Engineer

ABOUT THE ROLE:

We are looking for an accomplished Senior Data Engineer for the design, development, & deployment of cutting-edge data solutions. In this leadership role, you will drive technical strategy, mentor a team of data engineers, and collaborate with cross-functional stakeholders to deliver impactful, production-ready data systems. You will serve as technical authority on ingestion, processing and transformation, data modeling, translating complex business problems into scalable solutions.

KEY RESPONSIBILITIES:

  • Design & implement governed ingestion pipelines consuming on-site & off-site events into canonical event schema
  • Build and maintain Kafka or Kinesis-based streaming pipelines with schema validation, data quality checks, and alerting on source failures
  • Integrate batch ingestion from databases and alongside the streaming path, managing orchestration via Airflow or equivalent
  • Enforce privacy flags at ingestion time quarantining or anonymising events without valid consent before they reach downstream layers
  • Maintain a raw event store with partitions by date and source, serving as the audit source for the pipeline
  • Implement deterministic identity across fragmented systems
  • Design identity logic spanning on-site sessions, CRM records etc.
  • Monitor identity resolution quality — match rate, false positive rate, unresolved session ratio — and iterate on matching logic to improve coverage over time
  • Compute windowed aggregations over the event store
  • Build and maintain governed, versioned feature definitions in a feature store ensuring normalization, encoding, and embedding lookup logic is consistent between training and serving pipelines
  • Collaborate with data scientists to translate model feature requirements into production-grade pipeline implementations with no training-serving skew
  • Implement automated data contract validation between pipeline layers, schema compatibility checks, completeness assertions, and anomaly detection
  • Define and enforce SLAs on pipeline freshness, completeness, and accuracy with monitoring dashboards and escalation paths for breaches
  • Ensure PII is classified, tagged, and handled according to GDPR and CCPA requirements at every stage of the pipeline like ingestion, storage, feature compute, and serving
  • Contribute and maintain data lineage document, full traceability from raw event to feature to model prediction
  • Work closely with the Data Architect to implement data contracts and schema agreements between pipeline layers, and flag design risks early
  • Partner with ML Engineers to ensure feature pipelines meet model training and online inference requirements
  • Participate in code reviews, contribute to engineering standards, and mentor junior engineers where applicable

REQUIRED QUALIFICATIONS:

  • 6-8 years of hands-on data engineering experience, with at least 2 years in a senior or lead capacity on production systems
  • Demonstrable experience building and operating event-driven, streaming data pipelines at scale in a production environment
  • Prior experience on a personalization, recommendation, or user behavioral analytics platform is strongly preferred
  • Experience working within a regulated data environment like GDPR, CCPA, or equivalent.
  • Batch and structured streaming, including windowed aggregations, stateful processing, and performance tuning
  • Event streaming platforms including producer/consumer design, partition management, and exactly-once semantics
  • Data Lakehouse engineering like Delta Lake, Apache Iceberg, or equivalent, including ACID transactions, schema evolution, and time travel
  • Pipeline orchestration using Apache Airflow, AWS Glue, Azure Data Factory, or Databricks Workflows including DAG design, dependency management, and failure handling
  • SQL and Python at production standard quality with clean, tested, version-controlled code
  • Cloud data platform exposure on at least one of: AWS (S3, Glue, Kinesis, Redshift), Azure (ADLS, Data Factory, Synapse), or Databricks
  • Feature store design and operation: Databricks Feature Store, Feast, Tecton, or equivalent
  • Data quality frameworks: Great Expectations, dbt tests, or equivalent for automated pipeline validation
  • Deterministic and probabilistic matching, entity deduplication, graph-based stitching
  • CI/CD for data pipelines with automated testing, deployment, and monitoring using GitHub Actions, Azure DevOps, or equivalent

PREFERRED QUALIFICATIONS:

  • Experience with graph data modelling and graph processing frameworks — Neo4j, Amazon Neptune, GraphX, or GraphFrames
  • Familiarity with vector embedding pipelines - batch encoding of text at scale, embedding storage, and integration with ANN search infrastructure
  • Working knowledge of NLP pipeline engineering — tokenisation, embedding generation, and chunking for unstructured text at scale
  • Experience with real-time feature serving and integration like Redis, DynamoDB, or equivalent cache stores
  • Familiarity with MLflow or equivalent model registry — understanding of how feature pipelines connect to model training and deployment workflows
  • Experience with data mesh or federated data architecture patterns in multi-team environments
  • Strong written and verbal communication — able to explain complex pipeline design decisions to non-engineering stakeholders clearly
  • Able to make design decisions where requirements are incomplete and flag risks proactively
  • Collaborative working style with cross-functional teams spanning data engineering, ML, platform, and product
  • Understanding that interfaces between pipeline layers are as important as the implementations within them
  • Takes responsibility for pipeline reliability, DQ, and SLA adherence end-to-end, not just the code written

WHAT WE OFFER:

  • Awesome Culture: Creative Synergies has a flat organization and an agile culture of positivity, entrepreneurial spirit, customer centricity, celebrating technical excellence, teamwork, and meritocracy
  • Opportunity to work with Customers who are technology Leaders (including Global Fortune 500 Customers) & work on Real-World Problems that matter and are often mission-critical
  • Leadership role with significant influence over AI strategy and team direction.
  • Access to state-of-the‑art GPU infrastructure and cutting-edge AI tools.
  • Competitive compensation package with performance‑based incentives.
  • Flexible working arrangements with hybrid options.
  • Continuous learning budget for conferences, courses, and certifications.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Data Scientist
Senior Data Scientist

Confidential • Bengaluru

Hybrid
INR 2,600,000 - 4,200,000
Hybrid working
GPU infrastructure access
Continuous learning budget
+2
Senior Data Engineer (Data & Analytics)
Senior Data Engineer (Data & Analytics)

Summit Consulting Services • Ernakulam

On-site
INR 1,000,000 - 1,800,000
Collaborative culture
Opportunity to influence architecture
Modern cloud-native tech stack
Lead Data Engineer
Lead Data Engineer

PocketFM • Bengaluru

On-site
INR 3,000,000 - 5,400,000
Health insurance
Paid time off
Remote learning budget
Senior Data Engineer
Senior Data Engineer

DATAECONOMY Inc • Hyderabad

On-site
INR 1,500,000 - 2,100,000
Sr Data Engineer, Data Eng & Governance
Sr Data Engineer, Data Eng & Governance

Via Licensing Corporation • Bengaluru

On-site
INR 1,800,000 - 3,000,000
Sr Data Engineer, Data Eng & Governance
Sr Data Engineer, Data Eng & Governance

Dolby Laboratories • Bengaluru

On-site
INR 2,600,000 - 3,800,000
Sr Data Engineer, Data Eng & Governance
Sr Data Engineer, Data Eng & Governance

Dolby • Bengaluru

On-site
INR 4,000,000 - 6,000,000
Data Engineering Pipeline Engineer – Role Description
Data Engineering Pipeline Engineer – Role Description

Innoventes Technologies • Bengaluru

On-site
INR 1,000,000 - 1,500,000
Data Science Engineer
Data Science Engineer

NAVVYASA CONSULTING PRIVATE LIMITED • Gurugram District

On-site
INR 1,000,000 - 2,000,000
Opportunity to work with a fast-growing SaaS company
Dynamic environment impacting product growth
Career growth in advanced data science and AI
Technical Lead - Data Engineer (Data&AI)
Technical Lead - Data Engineer (Data&AI)

Srijan Technologies PVT LTD • Gurugram District

On-site
INR 4,000,000 - 7,500,000