Data Engineer

Xenonstack

Mohali, Chandigarh

On-site

INR 1,200,000 - 2,100,000

Full time

6 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Xenonstack is seeking a data engineer in Mohali to build and operate data pipelines powering AI workloads. The role focuses on batch and streaming workflows, ETL/ELT orchestration, and lakehouse ingestion using Python, SQL, Airflow, and Spark.

You will model data for analytics and ML, ensure end-to-end data quality, and optimize costs and performance while implementing governance. 2–4 years of experience required.

Qualifications

  • 2–4 years of professional data engineering experience.
  • Strong Python and advanced SQL with window functions and joins.
  • Hands-on with production data pipelines and orchestration tools.
  • Experience with distributed processing (Spark/PySpark) in cloud platforms.
  • Knowledge of data modelling and one cloud data platform (Databricks/Snowflake/BigQuery/Redshift).
  • Git, Docker and at least one cloud provider (AWS/Azure/GCP).
  • Familiarity with AI coding assistants and reviewing generated code.

Responsibilities

  • Build and operate batch and streaming data pipelines.
  • Design, develop, and maintain ETL/ELT workflows with orchestration tools.
  • Work with Spark on large datasets and manage lakehouse ingestion.
  • Model data for analytics and ML consumption, including feature prep.
  • Own data quality, tests, lineage, and monitoring end-to-end.
  • Optimize for cost and performance; manage access controls and governance.
  • Collaborate directly with AI engineering, platform, and product teams.

Skills

Python
Advanced SQL
Window functions
Query optimisation
Complex joins
Airflow
Dagster
Prefect
PySpark
Databricks
Snowflake
BigQuery
Redshift
Git
Docker
Cloud providers (AWS/Azure/GCP)
Data modelling
AI coding assistants

Education

BE / B.Tech in Computer Science or related field

Tools

Airflow
Dagster
Prefect
Spark (PySpark)
Databricks
Snowflake
BigQuery
Redshift
Kafka
Flink
Delta Lake / Iceberg / Hudi
dbt
Kubernetes
Terraform
CI/CD

Job description

Role & responsibilities

Build and operate the data pipelines that agentic AI systems run on.

  • Design, build, and maintain batch and streaming data pipelines in Python and SQL
  • Build and orchestrate ETL/ELT workflows using Airflow, Dagster, or similar
  • Work with distributed processing frameworks Spark or equivalent — on large datasets
  • Ingest data from APIs, databases, files, and event streams (Kafka or similar) into the lakehouse
  • Model data for analytics and for machine learning consumption, including feature preparation
  • Build and maintain the data layer behind retrieval-augmented generation — chunking, embedding pipelines, and vector store population
  • Own data quality end to end: validation, testing, monitoring, lineage, and alerting on pipeline failures
  • Optimise for cost and performance — query tuning, partitioning, storage format decisions
  • Implement access controls and governance appropriate to enterprise and regulated customers
  • Work directly with AI engineering, platform, and product rather than through handoffs

You will be building the data foundation for Agentic AI solutions and system, where pipelines feed live agent workflows, not just dashboards, and both correctness and latency matter.

Preferred candidate profile
  • 2–4 years of professional data engineering experience
  • Strong Python and advanced SQL — window functions, query optimisation, complex joins
  • Hands-on experience building production pipelines with an orchestration tool (Airflow, Dagster, Prefect)
  • Experience with a distributed processing framework, typically Spark (PySpark)
  • Working knowledge of a cloud data platform — Databricks, Snowflake, BigQuery, or Redshift
  • Comfortable with Git, Docker, and at least one cloud provider (AWS, Azure, or GCP)
  • Understanding of data modelling — dimensional modelling or one big table, and when each applies
  • Fluent with AI coding assistants in your daily workflow, and able to review what they produce rather than ship it unread
  • B.E. / B.Tech in Computer Science, Data Engineering, or a related field. Strong self-taught engineers with demonstrable production work are equally welcome
Good to have
  • Streaming experience — Kafka, Flink, Spark Structured Streaming
  • Lakehouse table formats — Delta Lake, Apache Iceberg, or Hudi
  • dbt for transformation and testing
  • Vector databases and embedding pipelines for RAG
  • Kubernetes, Terraform, or CI/CD for data infrastructure
  • Data observability and lineage tooling
  • Exposure to feature stores or ML pipeline tooling
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Engineer
Data Engineer

Drishti Consultants • Bengaluru

Hybrid
INR 1,200,000 - 2,400,000
Data Engineer
Data Engineer

Fruveggie Technology • Mumbai

On-site
INR 900,000 - 1,500,000
AI Data Engineer
AI Data Engineer

EXL • Gurugram District

On-site
INR 1,500,000 - 2,800,000
Lead Engineer - Data & AI
Lead Engineer - Data & AI

Quest Global • Bengaluru

On-site
INR 1,000,000 - 1,500,000
Senior Data Engineer (AI/ML)
Senior Data Engineer (AI/ML)

Neolatika • Maharashtra

On-site
INR 2,000,000 - 3,600,000
Technical Lead - Data Engineer (Data&AI)
Technical Lead - Data Engineer (Data&AI)

Srijan: Now Material • Gurugram District

On-site
INR 3,000,000 - 6,000,000
Data Engineering Pipeline Engineer – Role Description
Data Engineering Pipeline Engineer – Role Description

Innoventes Technologies • Bengaluru

On-site
INR 1,000,000 - 1,500,000
Data Engineer
Data Engineer

Advance Career Solutions • Pune District, Chennai District, Bengaluru

Hybrid
INR 1,200,000 - 2,800,000
Technical Lead - Data Engineer (Data&AI)
Technical Lead - Data Engineer (Data&AI)

Srijan Technologies PVT LTD • Gurugram District

On-site
INR 4,000,000 - 7,500,000
Data Engineer
Data Engineer

Exillar • Ahmedabad District

On-site
INR 600,000 - 1,200,000
Growth opportunities
Structured career path
Culture of continuous learning