Data Curation Engineer

Keka Technologies Private Limited

Bengaluru

On-site

INR 2,000,000 - 4,000,000

Full time

7 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Flam, an AI-native technology company in Bengaluru, seeks a data engineering specialist to build and maintain the data layer for our LLM/TTs/video pipelines. You will write Python daily, design scalable ingestion, cleaning, deduplication, labeling, and versioning of datasets used to train and evaluate models.

You’ll curate multilingual and code-mixed Indic data, ensure data quality, and manage audio/video assets.

Qualifications

  • Proficient Python for large-scale data processing.
  • Experience building data pipelines for ML/LLM workloads.
  • Strong understanding of data quality, cleaning, deduplication, and versioning.

Responsibilities

  • Build ingestion and cleaning pipelines at scale.
  • Generate and curate datasets for training/evaluation of models.
  • Handle multilingual and code-mixed Indic data and ensure data hygiene.
  • Own data provenance and versioning for reproducibility.

Skills

Strong Python
Data engineering fundamentals
ML data pipelines
Data quality

Tools

Parquet
WebDataset
HuggingFace Datasets
S3/GCS/R2

Job description

Flam is building the next generation of interactive media through its content format. We are an AI-native technology company transforming how brands and consumers interact through immersive, interactive content. Our technology enables rich, app-less experiences that can be launched instantly on smartphones, creating a fundamentally different way for brands to engage consumers. We are backed by leading investors and already work with some of the world's largest brands. We are now building Flicks, our interactive media format for the US market.

About the role :

Every model we ship at Flam is bounded by the data that trained it, and every eval number we report is only as trustworthy as the eval set behind it. This role owns both. You will build and run the data layer underneath our LLM Falcon, our TTS system Finesse). That means text, audio, and video, cleaned, deduplicated,filtered,labelled, synthetically generated where real data doesn't exist, and versioned so that six months from now we can say exactly what went into a checkpoint. This is an engineering role. You will write Python every day. It is not an annotation or labelling-management position, though you will design annotation guidelines and quality-check what comes back.

What you'll do

Build ingestion and cleaning pipelines at scale — deduplication (exact, near-dup, semantic), quality filtering with classifier-based scoring, PII stripping,language identification, format normalization.

Generate synthetic data where real data is scarce or expensive: LLM generated instruction and preference data, code-mixed Indic text, rendered scenes for image models, augmented audio for speech models.

Curate multilingual and code-mixed Indic corpora. A large part of our differentiation is Indic-language quality, and a large part of that comes down

to data hygiene most pipelines get wrong.

Build and maintain evaluation sets. Design them so they measure what we think they measure, keep them uncontaminated, and version them properly

Handle multimodal data - audio segmentation and transcript alignment for TTS/ASR, face and video preprocessing for avatar training, image-caption pair curation.

Own data provenance and versioning. Which files, which filters, which version, which run. This should be answerable in one command, not one

afternoon. Write annotation guidelines and audit annotation quality when human labelling is in the loop.

What we're looking for:

Strong Python. You are comfortable processing datasets far larger than memory, and you know when to reach for a database instead of a script.

Practical data engineering fundamentals — streaming, chunking, parallelism, checkpointing long jobs, handling malformed input without losing the run.

Familiarity with the modern data formats and tooling of ML Parquet,WebDataset, HuggingFace Datasets, object storage S3/GCS/R2.

Genuine care about data quality. You should find it uncomfortable when a dataset has duplicates in it.

Enough ML understanding to know how a data decision propagates into model behaviour — why dedup matters, why eval contamination invalidates a benchmark, what a bad filter does to a distribution's tails.

Strongly preferred:

Experience curating training data for LLMs, speech models, or image/video models specifically.

Working knowledge of embeddings and vector search for semantic deduplication and retrieval FAISS, Milvus, or similar).

Native or near-native fluency in one or more Indian languages, with the ability to judge quality in it — this is a real asset here, not a checkbox.

Audio or video processing experience (ffmpeg, torchaudio, forced alignment).

Synthetic data generation using LLMs, or 3D rendering pipelines Blender, Unreal) for visual data.

Not required:

Experience with our exact toolchain.

Why this role matters :

At most companies, data work is what gets handed to whoever is available. Here it's a named role with a named owner, because our model quality is directly downstream of it. You'll work alongside the engineers training the models, and you'll see your filtering decisions show up in benchmark numbers within the same quarter. If you want to move into model training over time, this is a strong path into it and we'll support that.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Artificial Intelligence Engineer
Artificial Intelligence Engineer

Flam • Bengaluru

On-site
INR 1,100,000 - 1,700,000
Data Engineer-Speech and language Data
Data Engineer-Speech and language Data

gnani.ai • Bengaluru

Hybrid
INR 1,500,000 - 2,100,000
Data Engineer - Data Platform
Data Engineer - Data Platform

Firmable • Kolkata District

On-site
INR 900,000 - 1,500,000
Senior AI/ML Engineer
Senior AI/ML Engineer

Dolby Laboratories • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Flex Work approach
Senior AI/ML Engineer
Senior AI/ML Engineer

Via Licensing Corporation • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Senior Manager, Data & AI Platform
Senior Manager, Data & AI Platform

Via Licensing Corporation • Bengaluru

On-site
INR 3,000,000 - 5,400,000
Flexible work policy
Competitive compensation
Benefits package
Senior AI/ML Engineer
Senior AI/ML Engineer

Dolby • Bengaluru

On-site
INR 2,800,000 - 6,000,000
Senior Engineer – Data
Senior Engineer – Data

Orbital • Hyderabad

On-site
INR 2,000,000 - 4,000,000
Member of Technical Staff, India
Member of Technical Staff, India

Human Archive • India

On-site
INR 2,000,000 - 3,400,000
Sr Software Engineer, Data & AI Platform
Sr Software Engineer, Data & AI Platform

Dolby Laboratories • Bengaluru

Hybrid
INR 1,800,000 - 2,500,000
Flexible work approach
Excellent compensation and benefits