Data Platform Engineer

BharatGen

Maharashtra

On-site

INR 1,000,000 - 1,500,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

A technology company focused on AI is seeking a Data Platform Engineer to design scalable platforms and robust data pipelines for Generative AI. You will work closely with researchers and ML engineers, ensuring high-quality dataset creation and operational reliability. Ideal candidates should have a degree in Computer Science or Data Engineering with experience in distributed systems and cloud platforms. This is a full-time position based in Maharashtra, India.

Qualifications

  • 3+ years of industry experience in data engineering or related field.
  • Exposure to data lifecycle management, including DataOps.

Responsibilities

  • Design and build scalable platforms for diverse datasets.
  • Develop robust data pipelines for Generative AI and LLM training.
  • Implement governance and observability for data quality.
  • Optimize platform performance and cost.
  • Collaborate closely with researchers and ML engineers.

Skills

Proficiency in distributed systems and frameworks (e.g., Kafka, Ray, PySpark)
Strong programming skills in Python, Scala, or Go
Experience with building and optimizing data pipelines
Expertise in data scraping and crawling frameworks
Experience with cloud platforms (AWS, GCP, Azure)
Deep understanding of data platform design
Familiarity with Infrastructure-as-Code tools
Expertise in managing unstructured data

Education

Bachelor’s or Master’s degree in Computer Science, Data Engineering

Tools

Docker
Kubernetes
Terraform
CloudFormation

Job description

Global Talent Acquisition Professional | APAC, UK & Ireland Hiring | Candidate Experience Champion | Stakeholder Management

Job Summary:

BharatGen is on a mission to create AI that truly represents the diversity, culture, and unique context of India. At the heart of this mission lies the need for robust, scalable infrastructure to build multilingual and multimodal datasets that power foundational AI models. We’re seeking a skilled Data Platform Engineer to build scalable tools, platforms, and pipelines tailored for processing large-scale, multilingual, multimodal datasets critical for foundational AI models.

In this role, you will build scalable data pipelines to ingest, transform, and prepare data from diverse sources—text, speech, images, and video—making it ready for Generative AI model training. Your work will involve developing and managing the underlying platform while addressing challenges like governance, security, observability, lineage, and scalability. The outcomes of your work will include efficient tools for data processing, a reliable data platform, and high-quality datasets tailored to the evolving needs of large-scale AI and LLM training.

Collaborating closely with researchers and ML engineers, you will play a pivotal role in enabling BharatGen to deliver state-of-the-art AI models, contributing to the advancement of India’s AI ecosystem through innovative data engineering solutions.

Key Responsibilities:

  • Design and Build Scalable Platforms: Develop distributed infrastructure for ingesting, processing, and transforming diverse datasets (text, speech, images, video) at terabyte to petabyte scale.
  • Develop Robust Data Pipelines: Create reliable, scalable pipelines to prepare datasets for Generative AI and LLM training.
  • Implement Governance and Observability: Build frameworks for data lineage, monitoring, and access control to ensure data quality and operational reliability.
  • Optimize Performance and Cost: Enhance platform performance and resource utilization using cost-effective strategies, including GPU-accelerated preprocessing.
  • Collaborate and Innovate: Work closely with researchers and ML engineers to adapt platforms and data pipelines to evolving LLM requirements, addressing various data challenges.
  • Drive Innovation: Stay updated on emerging tools, frameworks, and best practices to implement cutting-edge solutions for large-scale dataset creation.

Minimum Qualifications and Experience:

  • Bachelor’s or Master’s degree in Computer Science, Data Engineering, or a related field with 3+ years of industry experience.

Required Skills:

  • Proficiency in distributed systems and frameworks (e.g., Kafka, Ray, PySpark) for scalable data workflows.
  • Exposure to end-to-end data lifecycle management, including DataOps.
  • Strong programming skills in Python, Scala, or Go, with a focus on high-performance pipeline development.
  • Experience with building and optimizing data pipelines, including ETL processes, data modeling, and integration into scalable workflows.
  • Expertise in data scraping, crawling frameworks, and modern dataset development techniques such as synthetic data generation techniques.
  • Experience with cloud platforms (AWS, GCP, Azure) and container orchestration (Docker, Kubernetes).
  • Deep understanding of data platform design, including data architecture, metadata tracking, data lineage, observability, monitoring, and scalability best practices.
  • Familiarity with Infrastructure-as-Code tools (e.g., Terraform, CloudFormation), CI/CD pipelines, relational/NoSQL databases, and GPU-accelerated workflows.
  • Familiarity with visualization and monitoring tools for lifecycle management and pipeline performance tracking.
  • Expertise in managing unstructured data (text, speech, or multimodal datasets) for high-performance use cases, ideally in the context of LLM/AI datasets.
  • Understanding of challenges in scalable data engineering, including ingestion, transformation, and storage optimization for large-scale accelerated workflows.

Seniority level: Associate

Employment type: Full-time

Job function: Other

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Data Engineer
Data Engineer

Bsri Solutions • Chennai District, Bengaluru

On-site
INR 1,500,000 - 2,300,000
Senior Data Engineer (with AI/ML experience) India
Senior Data Engineer (with AI/ML experience) India

IDT Corporation • Bengaluru

On-site
INR 6,433,823 - 9,191,176
Competitive salary
Stability and growth opportunities
Referral program
+2
Senior Data Engineer – Knowledge Graph & AI Platform
Senior Data Engineer – Knowledge Graph & AI Platform

AiFA Labs • Hyderabad

On-site
INR 7,194,244 - 9,892,086
Competitive compensation with equity options
Flexible remote/hybrid work setup
Learning budget and conference support
Lead AI/ML Engineer
Lead AI/ML Engineer

Relanto • Bengaluru

On-site
INR 1,500,000 - 2,000,000
Senior Data Engineer
Senior Data Engineer

ACS International India Pvt. Ltd. (ACSII) • Maharashtra

On-site
INR 1,200,000 - 1,800,000
Associate Director - Data and AI
Associate Director - Data and AI

Welldoc,-Inc • Bengaluru

Hybrid
INR 4,000,000 - 7,000,000
Senior Data Science Manager
Senior Data Science Manager

Tredence Inc. • Pune District, Gurugram District, Bengaluru

On-site
INR 4,000,000 - 8,000,000
Senior Data Scientist
Senior Data Scientist

Ganit Business Solutions • Mumbai

On-site
INR 2,500,000 - 4,200,000
Senior Data Engineer (with AI/ML experience) India
Senior Data Engineer (with AI/ML experience) India

IDT • Pune District

On-site
INR 4,200,000 - 6,800,000
Remote work opportunity
Annual performance review
Career growth opportunities
+2
Associate Director - Data and AI
Associate Director - Data and AI

Welldoc • Bengaluru

Hybrid
INR 6,000,000 - 9,000,000