Big Data Engineer

Wingify

Gurugram District

On-site

INR 1,400,000 - 2,200,000

Full time

6 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Wingify is seeking a Big Data Engineer to design, build, and operate large-scale data pipelines and analytics infrastructure in a cloud-native environment. You will own end-to-end data ingestion, transformation, warehousing, and observability, with ClickHouse as the primary analytical store.

You will collaborate with data scientists, analysts, product engineers, and other teams to model data, optimize queries, and ensure reliability and cost-efficiency as data volumes grow.

Qualifications

  • Strong SQL and Python skills with production-grade code.
  • Hands-on experience with Airflow or comparable orchestration.
  • Hands-on experience with dbt or similar ELT tooling.
  • Expert-level ClickHouse production experience (MergeTree, partitions, materialized views).
  • Experience with a cloud-based data warehouse like BigQuery, including modeling and optimization.
  • Experience with NoSQL/Key-Value stores (Bigtable, DynamoDB, Redis).
  • Understanding of ETL/ELT and dimensional data modeling.
  • Experience with GCP and/or AWS and familiarity with Docker, Git, and CI/CD.
  • Focus on data quality, observability, and reliability.

Responsibilities

  • Design, build, and maintain robust batch and streaming data pipelines ingesting from multiple sources.
  • Build and operate Airflow DAGs with scheduling, retries, backfills, and idempotency.
  • Develop analytics-ready datasets using dbt with staged, intermediate, and mart layers.
  • Own ClickHouse as the primary analytical store: schema, partitions, materialized views, distributed/replicated tables, query/memory optimization.
  • Collaborate with data scientists, analysts, product engineers, and other teams to model data and optimize queries.
  • Implement monitoring, logging, and observability for pipelines and services; manage SLAs and incident response.

Skills

SQL
Python
Apache Airflow
dbt
ClickHouse
BigQuery / Cloud DW
NoSQL databases
GCP / AWS
Docker / CI/CD
Data quality & observability

Tools

ClickHouse
Airflow
dbt
BigQuery
Docker
Git
CI/CD
Kubernetes

Job description

Job Description:
About the Role

We are looking for a Big Data Engineer to design, build, and operate large-scale data pipelines and analytical infrastructure that transform high-volume raw data into reliable, query-ready datasets for analytics, reporting, and data-driven products.

Our data platform ingests and processes data from multiple sources and serves analytics, data science, product, and downstream applications. In this role, you will own data pipelines end-to-end—from ingestion and transformation to warehousing, orchestration, data quality, and observability.

A key part of the role is owning ClickHouse as our primary analytical data store. You will be responsible for designing scalable data models, optimizing query performance, and ensuring the platform remains reliable and cost-efficient as data volumes and workloads grow.

You will work closely with data scientists, analysts, product engineers, and other engineering teams to build a modern, cloud-native data platform.

What You'll Do
  • Design, build, and maintain robust batch and streaming data pipelines that ingest data from multiple sources into analytical data stores.
  • Build and operate Apache Airflow DAGs, including scheduling, dependencies, retries, backfills, idempotency, concurrency, and failure handling.
  • Develop analytics-ready datasets using dbt, following well-structured staging, intermediate, and mart layers with appropriate tests, documentation, and incremental models.
  • Own ClickHouse as the primary analytical store, including:
  • Schema and table design using the MergeTree family of engines
  • Partitioning and sorting/primary key strategies
  • Materialized views
  • Distributed and replicated table architectures
  • Query and memory optimization
  • High-volume data ingestion and performance tuning
  • Work with BigQuery where cloud data-warehouse patterns are appropriate, including data modeling and query/cost optimization.
  • Design and operate NoSQL and key-value data stores, including Bigtable, DynamoDB, and Redis, based on specific access patterns and performance requirements.
  • Build and maintain data-quality frameworks covering validation, testing, freshness, completeness, reconciliation, and anomaly detection.
  • Implement monitoring, alerting, structured logging, and observability for data pipelines and services.
  • Own pipeline SLAs, incident response, troubleshooting, and root-cause analysis.
  • Manage backfills, safe re-runs, schema evolution, and data migrations while minimizing downstream impact.
  • Build reproducible, containerized environments using Docker and contribute to CI/CD and Infrastructure as Code practices.
  • Partner with analysts, data scientists, product managers, and product engineers to translate business and technical requirements into scalable data models and pipelines.
  • Continuously improve pipeline reliability, scalability, performance, and infrastructure cost efficiency.
Must-Have Requirements
  • 4-6 years of experience in data engineering or a closely related field, with strong hands‑on production experience.
  • Expert-level SQL and strong Python skills, with experience writing production‑grade, maintainable, and well‑tested code.
  • Strong hands‑on experience with Apache Airflow or a comparable workflow orchestration platform, including DAG design, scheduling, retries, backfills, dependency management, idempotency, and concurrency.
  • Hands‑on experience with dbt or a comparable transformation/ELT framework, including modular models, testing, source management, documentation, and incremental processing.
  • Expert‑level production experience with ClickHouse. This is a core requirement and should include:
  • MergeTree engine family
  • Partitioning and primary/sorting keys
  • Materialized views
  • Distributed and replicated tables
  • Query and memory optimization
  • High‑volume ingestion and performance tuning
  • Production experience with a cloud‑based columnar/OLAP warehouse, such as BigQuery, including data modeling and performance/cost optimization.
  • Hands‑on experience with NoSQL and key‑value databases, such as Bigtable, DynamoDB, and Redis, including data modeling, partition/key design, access patterns, and caching strategies.
  • Strong understanding of ETL/ELT and dimensional/layered data modeling principles.
  • Experience with GCP and/or AWS and familiarity with Docker, Git, and CI/CD.
  • Strong focus on data quality, reliability, observability, and correctness.
  • Strong ownership, problem‑solving, and communication skills, with the ability to collaborate effectively across engineering, analytics, data science, and product teams.
Nice‑to‑Have
  • Experience with search platforms such as Elasticsearch, OpenSearch, Apache Solr, or Vespa, including indexing pipelines, schema design, and relevance/performance tuning.
  • Experience with Aerospike or other high‑performance, low‑latency distributed key‑value/NoSQL systems.
  • Experience building streaming and event‑driven pipelines using Kafka, Pub/Sub, or similar technologies.
  • Experience with Change Data Capture (CDC) patterns and technologies.
  • Experience with Apache Spark and data‑lake architectures using object storage such as GCS or S3.
  • Experience with Terraform or other Infrastructure as Code tools and Kubernetes.
  • Experience with data‑quality and observability tools such as Great Expectations, Soda, Monte Carlo, or advanced dbt testing.
  • Understanding of data platform cost optimization / FinOps practices.
  • Experience handling high‑volume e‑commerce, product catalog, behavioral, or event data.
  • Familiarity with a second backend programming language, particularly Go.
What Success Looks Like
  • Data pipelines are reliable, well‑tested, observable, and consistently meet freshness and completeness SLAs.
  • Data models are clean, scalable, documented, and trusted by analytics, data science, and product teams.
  • Data‑quality issues are identified before they impact downstream consumers.
  • Pipeline failures are diagnosed and resolved quickly, with clear root‑cause analysis and preventive actions.
  • ClickHouse and the broader analytical platform scale smoothly with growing data volumes and query workloads.
  • Data infrastructure remains performant and cost‑efficient as usage grows.
  • Downstream teams can confidently rely on the platform for analytics, reporting, experimentation, and data‑driven product experiences.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Data Engineer
Senior Data Engineer

AppSierra • India

On-site
INR 3,500,000 - 5,200,000
Senior Developer, Data Engineer
Senior Developer, Data Engineer

ICE • Hyderabad

On-site
INR 1,800,000 - 3,000,000
Principal Data Engineer
Principal Data Engineer

Dun & Bradstreet India • Hyderabad

On-site
INR 4,000,000 - 6,400,000
Software Engineer - Data & Scalability Platform (8-12 Yrs)
Software Engineer - Data & Scalability Platform (8-12 Yrs)

Cisco • Bengaluru

On-site
INR 3,000,000 - 6,000,000
Senior Data Engineer
Senior Data Engineer

DATAECONOMY Inc • Hyderabad

On-site
INR 1,500,000 - 2,100,000
Senior Developer, Data Engineer
Senior Developer, Data Engineer

Intercontinental Exchange Holdings, Inc. • Hyderabad

On-site
INR 3,000,000 - 5,500,000
Senior Developer, Data Engineer
Senior Developer, Data Engineer

ICE Clear Europe Limited • Hyderabad

On-site
INR 1,500,000 - 2,000,000
Data Engineer
Data Engineer

Allstate Solutions (ASPL) • Pune District, Bengaluru

Hybrid
INR 1,200,000 - 1,800,000
Data Engineer
Data Engineer

Staples India • Chennai District

On-site
INR 800,000 - 1,200,000
Data Engineer - ETL/Snowflake DB
Data Engineer - ETL/Snowflake DB

FirstHive | CDP+AI Data Platform • Bengaluru

On-site
INR 1,200,000 - 2,400,000