Data Engineer

ClearQuant

Ahmedabad District

On-site

INR 1,000,000 - 1,800,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

ClearQuant is seeking a hands-on Data Engineer to own the data backbone for our quantitative trading systems in India. You will build and maintain end-to-end data pipelines, ensuring clean, reliable, and reproducible data for research and execution.

The role focuses on data infra, not analytics or dashboards, with emphasis on Parquet/PostgreSQL/ClickHouse storage, data quality, and scalable pipelines. Linux, scheduling, and API/webscrape data integration are essential.

Qualifications

  • Bachelor’s or Master’s in CS/Engineering with 3–8 years of hands-on experience in Data Engineering / Data Infrastructure Roles.
  • Hands-on Python with pandas, NumPy, requests/httpx, file handling, logging, exception handling, retries, and rate-limit handling.
  • Proven experience building and maintaining production-grade data pipelines for time-series and messy real-world data.
  • PostgreSQL, ClickHouse (or similar OLAP systems), and Parquet, with robust data ingestion via APIs and webscraping.
  • Linux environments, debugging pipelines, logs, and scheduling (cron/systemd).
  • Knowledge of live vs historical data systems, latency, and data alignment across environments.
  • Mentoring junior team members and collaborating with the quant department to improve data reliability.

Responsibilities

  • Build and maintain data pipelines integrating data from broker APIs, exchanges, third-party providers, and internal datasets.
  • Clean and standardize data, handle corporate actions, symbol mapping, missing/inconsistent time series, and reconcile sources into unified schemas.
  • Design and manage data storage using Parquet, PostgreSQL, ClickHouse with efficient partitioning and indexing.
  • Develop data quality and validation systems to detect anomalies and ensure versioned, reproducible datasets.
  • Create an internal data access layer for live and historical data with consistent schemas for researchers.

Skills

Python
Pandas
NumPy
SQL
PostgreSQL
ClickHouse
Parquet
Linux
Data pipelines
ETL
APIs
PySpark
Kafka
Databricks
Web scraping

Education

Bachelor’s/Master’s in computer science, Engineering, or related field

Tools

PySpark
Kafka
Databricks
Snowflake
BigQuery
AWS Glue
Azure Data Factory

Job description

Type: Full-time Experience: 3 – 8 years

Application Form:

https://lnkd.in/dB-2a5wa

Qualifications: - Bachelor’s/ Master’s in computer science, Engineering, or related field with 3 – 8 years of hands-on experience in Data Engineering / Data Infrastructure Roles

About: - We are a proprietary quantitative trading firm across equities, derivatives, and crypto. This role owns the end-to-end data layer—from ingestion to delivery—ensuring all research and trading systems run on clean, reliable, and reproducible data. This is a core engineering role, responsible for the firm’s data backbone, not analytics or dashboards.

Primary Responsibility: You will be responsible for building and maintaining the data backbone of a quantitative trading system, ensuring that all downstream research and execution rely on high-quality, consistent data.

What This Role Is Not: This is not a cloud ETL, BI, dashboarding, Excel reporting, strategy development, or notebook-only role.

Experience with PySpark, Kafka, Databricks, Snowflake, BigQuery, AWS Glue, or Azure Data Factory is optional only, not the main requirement.

Key Responsibilities

  • Build and maintain data pipelines integrating data from broker APIs, exchanges, third party providers, and internal datasets, ensuring reliability and scheduling.
  • Clean and standardize data, handling corporate actions, symbol mapping, missing/inconsistent time series, and reconciling multiple sources into unified datasets with consistent schemas.
  • Design and manage data storage architecture using Parquet, PostgreSQL, ClickHouse, and CSV with efficient partitioning, indexing and querying.
  • Develop data quality and validation systems to detect anomalies, missing data, and duplicates; ensure versioned and reproducible datasets.
  • Build and maintain an internal data access layer and unified data layer for consistent retrieval of live and historical data with consistent schemas so that researchers can access data without directly querying databases for seamless strategy execution ensuring alignment between real-time and back testing datasets.
  • Manage changes in upstream data sources (APIs, websites, schemas) to ensure pipeline continuity and minimal disruption.
  • Ensure pipeline reliability on Linux systems with job scheduling, monitoring, logging, debugging failures, and maintaining fault-tolerant, observable, and recoverable systems.
  • Enable robust data backfilling and reprocessing for regenerating datasets when source data or processing logic changes.
  • Guide and review work of data analyst interns, ensuring structured, consistent, and integrated data outputs.
  • Work closely with quant department to streamline data access, troubleshoot issues, and improve data reliability, enabling faster and more accurate strategy development.

Requirements

  • Strong hands-on Python expertise with pandas, NumPy, requests/httpx, file handling, logging, exception handling, retries, and rate-limit handling.
  • Proven experience building and maintaining production-grade data pipelines, especially for time-series and messy real-world data (missing values, schema inconsistencies).
  • Hands-on experience with PostgreSQL, ClickHouse (or similar OLAP systems),and Parquet, along with building robust data ingestion pipelines via APIs and webscraping.
  • Comfortable in Linux environments, with the ability to debug pipelines, manage logs, and handle scheduling (cron/systemd).
  • Clear understanding of live vs historical data systems, including latency, consistency, and alignment across both environments.
  • Proven ability to design reliable, reproducible, and high-quality data systems, with experience of enforcing data standards and mentoring junior team members.
  • Knowledge of Indian/US Stock Markets is beneficial but not necessary.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Quantitative Data Engineer
Senior Quantitative Data Engineer

Qode Advisors Llp • Mumbai

On-site
INR 1,500,000 - 2,000,000
High ownership and autonomy
Exposure to quantitative research
Fast-paced work environment
Senior Quantitative Data Engineer
Senior Quantitative Data Engineer

Qode • Mumbai

On-site
INR 1,500,000 - 2,500,000
High ownership and autonomy
Opportunity to mentor others
Fast-paced environment with direct impact
Senior Data Engineer
Senior Data Engineer

DATAECONOMY Inc • Hyderabad

On-site
INR 1,500,000 - 2,100,000
Data Engineer - SQL+Python
Data Engineer - SQL+Python

Luxoft • Pune District

On-site
INR 2,000,000 - 4,000,000
Senior Data Engineer
Senior Data Engineer

Luxoft India • India

On-site
INR 1,500,000 - 2,000,000
Data Analytics Engineer
Data Analytics Engineer

EXL • Hyderabad, Pune District, Bengaluru

Hybrid
INR 1,200,000 - 1,800,000
Data Engineer (Contract)
Data Engineer (Contract)

NexTurn Inc. • Hyderabad

On-site
INR 1,200,000 - 1,800,000
Data Engineer
Data Engineer

Navikenz India • Bengaluru

On-site
INR 1,200,000 - 1,800,000
Data Engineer
Data Engineer

Deservely Technologies Pvt Ltd • Hyderabad

On-site
INR 1,500,000 - 2,300,000
Data Engineer
Data Engineer

ConveGenius.AI • Chennai District

On-site
INR 2,100,000 - 3,200,000