Senior Data Engineer / SSE

apna

Bengaluru

On-site

INR 2,800,000 - 4,500,000

Full time

11 days ago
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Apna is seeking a Senior Data Engineer to design, build, and operate the core data platform. You will work on large-scale data pipelines, lakehouse architectures, and orchestration systems that power analytics, product intelligence, and ML-driven decisions.

You will collaborate with product, analytics, ML, and backend teams to ensure data quality, reliability, and scalable data models for dashboards and insights.

Qualifications

  • Strong experience in data engineering at scale.

Responsibilities

  • Designing, building, and operating key parts of Apna's data platform.
  • Building scalable batch and near-real-time data pipelines.
  • Designing lakehouse architecture with Apache Hudi.
  • Working with query engines like Presto/Trino for analytical workloads.
  • Building and maintaining orchestration workflows with Apache Airflow.
  • Mentoring data engineers and influencing architecture decisions.

Skills

Data engineering
Airflow
Presto/Trino
Hudi
SQL
Python/Java/Scala
Data modeling
ETL/ELT
Data warehousing
Debugging/production ownership

Tools

Apache Airflow
Apache Hudi
Presto/Trino
Kafka
Spark
Hive

Job description

Role: Senior Data Engineer / SSE

Requirement: 1

Team: Data Platform / Engineering

Location: Work from Office - Domlur, Bangalore (5 days / week)

Experience : 4-6 Years of Experience

Why Join Apna

At Apna, data is central to how we build products, understand users, improve employer outcomes, power recommendations, and scale decision-making. This role gives you the opportunity to build the backbone of Apna's data platform and influence how data is used across the company.

You will work on real-world, high-scale problems across jobs, users, employers, communities, matching, growth, and AI-driven systems.

About The Role

Apna is looking for a Senior Software Engineer to build and scale our core data platform. This role will work on large-scale data pipelines, lakehouse architecture, query platforms, workflow orchestration, and data reliability systems that power analytics, product intelligence, machine learning, business dashboards, experimentation, and operational decision-making across Apna.

We are looking for someone who can think deeply about data architecture, design reliable pipelines, improve data quality, and help build a platform that can scale with Apna's growth.

Requirements
What You'll Own

You will be responsible for designing, building, and operating critical parts of Apna's data platform, including:

  • Building scalable batch and near-real-time data pipelines across product, business, growth, and ML use cases
  • Designing and improving our lakehouse architecture using technologies likeApache Hudi
  • Working with query engines such asPresto / Trinofor large-scale analytical workloads
  • Building and maintaining orchestration workflows usingApache Airflow
  • Creating reusable data models, curated datasets, and reliable data marts for analytics and product teams
  • Improving data platform reliability, observability, SLA tracking, lineage, and data quality checks
  • Optimizing storage, compute, query performance, and pipeline costs
  • Partnering with product, analytics, ML, and backend engineering teams to understand data needs and convert them into scalable platform solutions
  • Driving engineering standards around data modeling, schema evolution, partitioning, deduplication, backfills, replayability, and pipeline ownership
  • Mentoring data engineers and influencing architecture decisions across teams
What We're Looking For
Must Have
  • Strong experience indata engineering, preferably at scale
  • Hands-on experience withApache Airflowor similar orchestration systems
  • Strong knowledge ofPresto / Trinoor other distributed query engines
  • Good understanding ofApache Hudiconcepts such as:
    • Copy-on-write vs merge-on-read
    • Upserts and deletes
    • Incremental reads
    • Compaction
    • Clustering
    • Timeline and commits
    • Schema evolution
    • Partitioning strategy
  • Strong knowledge of distributed data processing and storage systems
  • Ability to design and build reliable ETL / ELT pipelines
  • Strong SQL skills and ability to debug complex data issues
  • Good understanding of different data architectures, including:
    • Data warehouse
    • Data lake
    • Lakehouse
    • Lambda architecture
    • Kappa architecture
    • Medallion architecture
    • Event-driven data architecture
  • Experience with data modeling for analytics and reporting
  • Strong programming skills in at least one language such asPython, Java, or Scala
  • Ability to reason about trade-offs between freshness, cost, reliability, latency, and complexity
  • Strong debugging and production ownership mindset
Good to Have
  • Experience with Kafka, Spark, Flink, Hive, Iceberg, Delta Lake, or BigQuery
  • Experience building internal data platforms or self-serve data infrastructure
  • Experience with data quality frameworks such as Great Expectations, Deequ, Soda, or custom validation systems
  • Exposure to ML feature pipelines or feature stores
  • Experience with metadata management, data catalogs, lineage, and governance
  • Experience with cloud infrastructure such as AWS, GCP, or Azure
  • Understanding of privacy, compliance, PII handling, and access control in data systems
What Success Looks Like
  • Critical business and product datasets are reliable, discoverable, and trusted
  • Pipelines are observable, recoverable, and have clear SLAs
  • Query performance improves across major analytical workloads
  • Data freshness and quality issues reduce significantly
  • Teams can build on top of the data platform faster without reinventing pipelines
  • The platform can scale with Apna's user, job, employer, and engagement data
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Data Engineer
Senior Data Engineer

Aon Corporation • Bengaluru

Hybrid
INR 2,200,000 - 4,200,000
Cab facility
Senior Data Engineer
Senior Data Engineer

INFOSLAB CONSULTANCY PRIVATE LIMITED • Hyderabad

On-site
INR 1,200,000 - 1,800,000
Data Engineer Pune · Hybrid Engineering · Full-time →
Data Engineer Pune · Hybrid Engineering · Full-time →

Woodfrog Tech OPC Private Limited • Pune District

On-site
INR 1,800,000 - 2,400,000
Hybrid working
Senior Data Engineer
Senior Data Engineer

Vidpro Consultancy Services • Pune District, Chennai District, Bengaluru

On-site
INR 1,800,000 - 2,800,000
Senior Data Engineer
Senior Data Engineer

Weekday 1 • Bengaluru

On-site
INR 8,000,000 - 20,000,000
Senior Data Engineer
Senior Data Engineer

Luxoft India • India

On-site
INR 1,500,000 - 2,000,000
Senior Data Engineer (Senior Individual Contributor / Technical Lead)
Senior Data Engineer (Senior Individual Contributor / Technical Lead)

Bootminds • Bengaluru

On-site
INR 2,000,000 - 2,600,000
Senior Data Engineer
Senior Data Engineer

Takeda • Bengaluru

On-site
INR 1,200,000 - 2,800,000
Senior Data Engineer - Python/SQL
Senior Data Engineer - Python/SQL

India Fan Corporation • Bengaluru

On-site
INR 1,800,000 - 2,400,000
Senior Data Engineer (Data & Analytics)
Senior Data Engineer (Data & Analytics)

Summit Consulting Services • Ernakulam

On-site
INR 1,000,000 - 1,800,000
Collaborative culture
Opportunity to influence architecture
Modern cloud-native tech stack