Data Engineer

CloudCrane

India

On-site

INR 1,400,000 - 2,000,000

Full time

29 hours ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

CloudCrane in India seeks a data engineer to own the data import, type inference, deduplication, and batch enrichment pipelines. You will also manage the embedding and indexing path and ensure data tests validate output correctness.

The role requires 3+ years building scheduled data pipelines, strong PostgreSQL skills, and proficiency in Python or TypeScript. Experience with messy real-world files and a focus on idempotent, resumable workflows is essential.

Qualifications

  • Three or more years building data pipelines that run on a schedule.
  • Strong SQL and a working understanding of how Postgres behaves under load.
  • Python or TypeScript, and the judgement to know when a transformation belongs in the database.
  • Handled genuinely messy real-world files and can describe encoding, nulls and duplicates.
  • Design for reruns: idempotency and resumability are instincts, not afterthoughts.

Responsibilities

  • Own the import path: CSV, TSV, JSONL and spreadsheets with limits enforced on actual reads.
  • Infer and confirm column types, identity columns and duplicates for correct downstream sorting and deduplication.
  • Build and tune batch enrichment with resumable runs and idempotent writes.
  • Own the embedding pipeline: what gets embedded and when it is recomputed.
  • Write data tests ensuring outputs match expected results, not just that the pipeline ran.

Skills

SQL
Python
TypeScript
Data pipelines
Deduplication

Tools

PostgreSQL
Vector stores

Job description

Own how a customer's data gets in, gets shaped, and stays correct as it moves through enrichment and into a release.

Everything the product promises depends on the data being right before anything clever happens to it. That is this role: the import path, the type inference, the identity column and the deduplication, the batch enrichment runs, and the immutable release snapshot that a live agent reads from.

The work is equal parts pipeline and correctness. A CSV with a decimal comma, a spreadsheet with merged cells, a file that claims to be 2MB and inflates to 2GB: each of these has already caused a real bug here, and each was fixed by making the pipeline refuse rather than guess.

You would also own the embedding and indexing path: what text represents a record, when it needs recomputing, and how the cache stays honest when a contract changes underneath it.

What you'd do
  • Own the import path: CSV, TSV, JSONL and spreadsheets, with limits enforced on what is actually read rather than what a header claims
  • Infer and confirm column types, identity columns and duplicates, so numeric sorting and dedup are correct downstream
  • Build and tune batch enrichment: resumable runs, idempotent writes, and caches that never serve a stale answer
  • Own the embedding pipeline: what gets embedded, when it is recomputed, and how it is stored
  • Write the data tests: not just that the pipeline ran, but that what came out is what should have
What we need
  • Three or more years building data pipelines that ran on a schedule and had someone depending on them
  • Strong SQL and a working understanding of how Postgres behaves under load
  • Python or TypeScript, and the judgement to know when a transformation belongs in the database
  • You have handled genuinely messy real-world files and can describe what you did about encoding, nulls and duplicates
  • You design for reruns: idempotency and resumability are instincts, not afterthoughts
Nice to have, not required
  • Experience with vector stores, embeddings or search indexing
  • Worked with a queue, a scheduler or a workflow engine in production
  • Data quality or observability tooling you built rather than bought
How we work
  • Small team, thin slices, shipped weekly. Nothing sits on a branch for a month.
  • Reviews are real. Every feature ships with its tests, and a bug found in review is cheaper than one found by a customer.
  • Safety-critical means the boring answer usually wins: fail closed, keep the receipt, don't guess.
  • We only claim what ships. That applies to the product, the roadmap and the offer letter.
How hiring works
  • 1 A 30-minute call about something you've built and what was hard about it.
  • 2 A working session on a real problem from this codebase. You can drive, or we can pair.
  • 3 A conversation about how you make decisions when the evidence is thin.
  • 4 References, then an offer. We aim to answer within a week at every stage.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Data Engineer
Senior Data Engineer

Commotion • India

On-site
INR 4,000,000 - 7,000,000
Software Engineer, Backend
Software Engineer, Backend

CloudCrane • India

On-site
INR 1,200,000 - 2,200,000
Tech Lead at Saleshandy
Tech Lead at Saleshandy

Ikigai Infotech LLP • Ahmedabad District

On-site
INR 4,000,000 - 6,500,000
Data Engineer
Data Engineer

Mechademy • Gurugram District

Hybrid
INR 2,000,000 - 4,000,000
Forward Deployed Engineer - Data
Forward Deployed Engineer - Data

Commotion • India

On-site
INR 1,200,000 - 2,400,000
Data Engineer - ETL/Snowflake DB
Data Engineer - ETL/Snowflake DB

FirstHive | CDP+AI Data Platform • Bengaluru

On-site
INR 1,200,000 - 2,400,000
Software Engineer, Business Systems
Software Engineer, Business Systems

Engg • Pune District

On-site
INR 900,000 - 1,400,000
Senior Data Engineer (Data Platform)
Senior Data Engineer (Data Platform)

Mico • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Founding Data Pipeline Engineer
Founding Data Pipeline Engineer

HuntingCube • Gurugram District

On-site
INR 1,200,000 - 2,500,000
Data Engineer
Data Engineer

Coretek • Kondapur

On-site
INR 1,200,000 - 1,800,000