Data Science Engineer

J&F

Chennai District

On-site

INR 1,200,000 - 1,800,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

J&F is seeking a Data Engineer / ML Data Pipeline Engineer to build and operate the data backbone of the enterprise AI platform. You will own ingestion, ETL/ELT workflows, and AWS-based architecture across raw to curated data.

You will implement data validation, observability, and dashboards to monitor quality, cost, and performance. Responsibilities include ML data workflow support, API development for validation runs, and end-to-end pipeline ownership with a focus on scalable, robust data

Qualifications

  • Python — production-grade scripting for file/JSON/schema processing and robust error handling.
  • SQL — strong hands-on skills incl. GROUP BY/HAVING, window functions, daily aggregates.
  • AWS S3 data handling — structuring buckets for raw/staging/curated data with versioning.
  • Data validation — set comparisons, duplicate/missing detection, structured pass/fail reporting.
  • ETL/ELT pipeline design — end-to-end ownership from source to business outcomes.
  • Query/warehouse judgment — know when to use Athena vs. Redshift/Snowflake with partitioning.

Responsibilities

  • Ingestion & ETL/ELT pipelines for heterogeneous project folders (PDFs, SVGs, IFC models, JSON, Excel).
  • AWS-based data architecture: S3 structuring, partitioning, versioning, lifecycle management; querying via Athena/Glue; warehousing as needed.
  • Data validation frameworks: GUID cross-referencing, schema enforcement, duplicate/orphan checks, structured reports.
  • Agent run logging & observability: design DB schema and pipelines to track AI agent runs.
  • AI Factory monitoring dashboards: operational and business dashboards for Power BI/QuickSight.
  • ML data pipeline support: dataset preparation, labeling workflows, human-in-the-loop review, and dataset versioning.
  • APIs: building FastAPI/Flask endpoints to trigger validation runs and expose processing status.
  • Data quality & testing: idempotent pipelines, quarantine/reject handling, regression testing, root-cause debugging.

Skills

Python
SQL
AWS S3
Data validation
ETL/ELT pipelines
Query/warehouse judgement

Tools

AWS Glue
Athena
Redshift
Snowflake

Job description

Role Summary

We Are Hiring a Data Engineer / ML Data Pipeline Engineer To Build And Operate The Data Backbone Of The Enterprise AI Platform

What You'll Own
  • Ingestion & ETL/ELT pipelines for heterogeneous project folders (PDF drawings, SVG files, IFC models, BBS.json bar-bending-schedule data, Excel exports, and AI agent output JSON).
  • AWS-based data architecture: S3 raw/staging/curated/outputs structuring, partitioning, versioning, and lifecycle management; querying via Athena/Glue and warehousing via Redshift or Snowflake as needed.
  • Data validation frameworks: GUID cross-referencing between SVG and BBS data, schema enforcement, duplicate/orphan detection, reference integrity checks, and structured validation reporting.
  • Agent run logging & observability: designing the database schema and pipelines that track every AI agent run (inputs, outputs, status, errors, cost, retries, reviewer feedback).
  • AI Factory monitoring dashboards: operational dashboards (failure rates, retries, latency, data quality) and business dashboards (throughput, cost per run, rework rate) for Power BI/QuickSight or equivalent.
  • ML data pipeline support: dataset preparation, labeling/annotation workflows, human-in-the-loop review tooling, and dataset versioning for models that classify or QC drawing issues.
  • APIs: designing and building FastAPI/Flask endpoints to trigger validation runs and expose agent processing status to internal tools.
  • Data quality & testing discipline: idempotent pipelines, quarantine/reject handling, regression and reconciliation testing, and root-cause debugging when pipelines or query performance degrade in production.
Key Skills — Non-Negotiable (Must-Have, Strong Level)
  • Python — production-grade scripting: file/folder handling, JSON/schema processing, clean error handling, not just notebook-level scripting.
  • SQL — strong hands-on ability, including GROUP BY/HAVING for duplicate detection, window functions, and daily aggregate/rate calculations (e.g., success-rate queries).
  • AWS S3 data handling — practical experience structuring buckets for raw/staging/curated data, versioning, and avoiding overwrite issues at scale.
  • Data validation — demonstrable experience building validation logic (set comparisons, duplicate/missing detection, structured pass/fail reporting), not just "I write assertions."
  • ETL/ELT pipeline design — end-to-end ownership of at least one pipeline: source → transform → storage → validation → monitoring → business outcome, with clear articulation of what they personally built.
  • Query/warehouse engine judgment — working knowledge of when to use Athena vs. Redshift vs. Snowflake (or equivalent), partitioning, clustering, sort/distribution keys, and storage format trade-offs (Parquet vs. JSON vs. CSV).
Key Skills — Good to Have
  • Dashboarding — Power BI / QuickSight (or equivalent) fact/dimension table design, KPI cards, drill-downs; medium-to-strong level is a plus but trainable.
  • FastAPI / Flask — building real endpoints with request/response schemas and basic error handling; especially valuable for validation-trigger and agent-status APIs.
  • ML data pipeline experience — dataset labeling, annotation platform design, train/test/validation splitting, dataset versioning; strong on the pipeline/data side rather than model training itself.
  • Human-in-the-loop / review tooling — experience building or contributing to browser-based labeling/review platforms (session persistence, label schema, export formats).
  • Large-scale metadata querying — experience making file discovery fast across large volumes (1,000+ projects, thousands of files each) via metadata index tables, event-based ingestion, or catalog tools like AWS Glue.

Skills:- Generative AI, LangGraph, ETL, databricks, Retrieval Augmented Generation (RAG) and FastAPI

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Engineer
AI Engineer

DocuSign, Inc. • Bengaluru

Hybrid
INR 1,500,000 - 2,500,000
Lead Engineer - Data & AI
Lead Engineer - Data & AI

Quest Global • Bengaluru

On-site
INR 1,000,000 - 1,500,000
Data Engineer (Databricks + AI)
Data Engineer (Databricks + AI)

Durapid Technologies Pvt Ltd • Bengaluru Urban

On-site
INR 1,500,000 - 2,300,000
VP, Process Improvement
VP, Process Improvement

Jobtailor • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Senior Data Scientist
Senior Data Scientist

Neurealm • Gurugram District

On-site
INR 900,000 - 1,300,000
Data Scientist
Data Scientist

Cognizant • Bengaluru Urban

On-site
INR 4,000,000 - 7,000,000
Delivery Lead-Data Engineer
Delivery Lead-Data Engineer

Acuity Analytics • Gurugram District

On-site
INR 1,500,000 - 2,500,000
Data Engineer - Lead
Data Engineer - Lead

Iris Software • Dadri

On-site
INR 1,500,000 - 2,500,000
Lead Data Engineer
Lead Data Engineer

PocketFM • Bengaluru

On-site
INR 3,000,000 - 5,400,000
Health insurance
Paid time off
Remote learning budget
Technical Lead - Data Engineer (Data&AI)
Technical Lead - Data Engineer (Data&AI)

Srijan Technologies PVT LTD • Gurugram District

On-site
INR 4,000,000 - 7,500,000