Data Scientist / Lead Data Engineer

GoComet

Bengaluru

On-site

INR 4,000,000 - 7,000,000

Full time

6 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

GoComet is seeking a senior data platforms engineer to own the end-to-end event pipeline and data quality. You will decide on architecture, warehouse modeling, and data freshness targets, ensuring reliable, reproducible metrics for customers.

You will lead, mentor, and own backfills and complex data transformations, with a strong emphasis on statistical rigour and operational excellence. You will design and defend data models, implement robust ETL with Airflow/Dagster/Prefect/Temporal, and

Qualifications

  • 5+ years building and operating production data platforms with 2+ years as architecture owner.
  • Designed a warehouse from first principles with dimensional models and MI/MD definitions.
  • Deep SQL and Python with advanced windowing, safe backfills, and idempotent transforms.
  • Own a production orchestrator with retries, SLAs, dependencies, and backfills.
  • Defensible stance on table formats and streaming vs change data capture decisions.
  • Demonstrated statistical rigour: experiment design, power calculations, appropriate tests.
  • Led engineers, performed hiring, reviews, and on-call pager duties.

Responsibilities

  • Own the event pipeline: late arrivals, restatements, and ingest vs event time gaps.
  • Define how restatements propagate to customers who already saw the data.
  • Make the serving path agent-grade with explicit freshness targets.
  • Enforce data quality as failing checks, not dashboards.
  • Plan backfills across 30+ countries with cost considerations and blast radius.
  • Applying statistical judgement before decisions and communicating limits clearly.

Skills

5+ years data platforms
Architecture ownership
Warehouse design
SQL
Python
Production orchestrator
Airflow
Dagster
Prefect
Temporal
Table formats reasoning
Statistics & experiments
Team leadership

Education

Bachelor's degree in a quantitative field

Tools

Airflow
Dagster
Prefect
Temporal

Job description

The Role

Every number a customer argues about.


Architecture is still open and yours to decide and defend: warehouse shape, table format, orchestration, where the streaming boundary sits against change capture or a tighter batch cadence, and how datasets are tiered with real service levels behind them.


This is not an Analytics Engineer role. They decide what a term means; you decide whether the number is computed correctly, arrives on time, and can bear the inference drawn from it.


If a metric means two different things in two places, that is theirs. If it means the right thing but is stale or irreproducible, it is yours.


What You Would Actually Be Doing

Own the Event Pipeline

Own the event pipeline behind shipment milestones: late arrivals, corrections that restate a milestone recorded three days ago, and the gap between event time and ingest time.


Define how a restatement propagates to something a customer has already seen.


Build for Agent-Grade Serving

Make the serving path good enough for an agent to depend on — with a stated freshness target per dataset and a read path whose p99 you know.


Because when a run stalls waiting for a number, a workflow with money attached stalls too.


Make Data Quality Fail, Not Warn

Build data quality as failing checks, not dashboards:



  • Null rates

  • Referential integrity breaks

  • Duplicate milestones

  • Distribution shifts when a partner changes a field


A check that only warns is a check nobody reads.


Enforce tenant isolation at the data layer and assert it through tests.


Contract rates are commercially sensitive between competitors who may share the same forwarders.


Own Expensive Backfills

Plan backfills that cost real money.


A reprocess across 30+ countries of historical data has a bill and a blast radius. You own both.


Apply Statistical Judgement

Bring statistical rigour to decisions.


Know what you can and cannot claim from observational data — and say so before the decision, not after it.


What We Look For

Must-Haves

Years are a floor, not the bar.


If you miss one line but are strong on the rest, apply and tell us which one.



  • 5+ years building and operating production data platforms, including at least 2 years as the person accountable for architecture others depended on.

  • Designed a warehouse from first principles — grain, dimensional models, slowly changing dimensions, additive vs. non-additive measures — and can describe a modelling trap you found in someone else’s work and fixed.

  • Deep SQL and Python: window functions, incremental and idempotent transformations, safe-to-rerun backfills, and pipelines that survive late-arriving data.

  • Hands-on ownership of a production orchestrator — Airflow, Dagster, Prefect, Temporal, or equivalent — including retries, service levels, dependency modelling, and backfill strategy.

  • A defensible position on table formats and streaming vs. change capture vs. tighter batch cadence, with the reasoning rather than the fashion.

  • Genuine statistical rigour: can design an experiment, defend the power calculation, and pick the right test for skewed or correlated data.

  • Has led engineers — hiring, reviews, mentoring — and carried a pager for datasets other teams depended on.


Also Good — None Required


  • You have run a data on-call rotation and can describe what you changed to make it quieter.

  • You have worked in a domain where being wrong about a time or quantity had a cost someone could name.


By Month Six

What Good Looks Like

We would rather tell you now what we would be measuring, so you can decide whether this is the job you want.


01 — Dataset Reliability

Every dataset a workflow depends on has a stated freshness target and a failing check behind it.


Tenant isolation in the data layer is asserted by a test that runs on every change.


03 — Team & On-Call Impact

You have hired or levelled up at least one engineer, and the data on-call is quieter than when you arrived.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Sr. Staff Backend Engineer
Sr. Staff Backend Engineer

PHIZENIX • Hyderabad

Hybrid
INR 4,000,000 - 7,000,000
Data Engineer - ETL/Snowflake DB
Data Engineer - ETL/Snowflake DB

FirstHive | CDP+AI Data Platform • Bengaluru

On-site
INR 1,200,000 - 2,400,000
Software Engineer III Chennai, India · On-site
Software Engineer III Chennai, India · On-site

Arcadia Power, Inc. • Chennai District

Hybrid
INR 2,000,000 - 4,000,000
Stock options
Hybrid work in Chennai
Medical insurance (self + 5 family)
+4
Senior Data Engineer
Senior Data Engineer

Minfy • India

On-site
INR 4,000,000 - 6,000,000
Senior Data Engineer
Senior Data Engineer

Minfy • Gurugram District

On-site
INR 4,500,000 - 6,000,000
Software Engineer III
Software Engineer III

Arcadia Power, Inc. • Chennai District

Hybrid
INR 2,800,000 - 5,600,000
Employee stock options
Hybrid work model
Medical insurance (self + family)
+2
Senior Business Analyst
Senior Business Analyst

Finkraft • Bengaluru

On-site
INR 2,000,000 - 3,000,000
Software Engineer II, Data Engineering
Software Engineer II, Data Engineering

Jobtailor • Bengaluru

On-site
INR 2,000,000 - 4,000,000
Fullstack Architect
Fullstack Architect

C5i • Bengaluru

On-site
INR 3,500,000 - 6,500,000
Senior Data Engineer
Senior Data Engineer

AppSierra • India

On-site
INR 3,500,000 - 5,200,000