Data Enginner

Stock Bazaar

Delhi

On-site

INR 1,200,000 - 2,200,000

Full time

26 hours ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Stock Bazaar in Delhi is seeking a data engineer to design, build and maintain real-time and batch data pipelines for market data ingestion. You will lead PySpark transformations for large datasets and ensure fault-tolerant processing.

Design and maintain the data lake architecture, implement Kafka streaming, and manage schema evolution. Docker/Kubernetes and AWS experience help you own CI/CD for data jobs and collaborate with DevOps.

Qualifications

  • 3+ years in data engineering or a closely related role.
  • Apache Kafka — production experience with streaming pipelines.
  • PySpark — batch and streaming, with tuning knowledge.
  • Data lake architecture — partitioning, formats (Parquet), query performance.
  • DevOps practices — Docker, Kubernetes, CI/CD.
  • Strong Python and SQL.
  • Cloud experience, preferably AWS (S3, EMR, Glue).
  • Bachelor's degree in Computer Science, Information Technology or related, or equivalent practical experience

Responsibilities

  • Design, build and maintain real-time and batch data pipelines for market data ingestion.
  • Build streaming pipelines using Apache Kafka — topics, partitioning, consumer groups, schema management.
  • Write distributed processing jobs in PySpark for large-scale transformation and aggregation.
  • Ensure pipelines are fault-tolerant, idempotent and recover cleanly from failures.
  • Design and maintain the data lake, including partitioning strategy, file formats and retention.
  • Model data for analytical querying and low-latency platform reads.
  • Manage schema evolution without breaking downstream consumers.
  • Build validation and reconciliation checks so bad or missing market data is caught before it reaches subscribers.
  • Monitor pipeline health, set up alerting, and own incident resolution and root-cause analysis.
  • Own deployment of data services using Docker and Kubernetes; CI/CD pipelines.
  • Manage cloud infrastructure (AWS) for data workloads, with an eye on cost.
  • Work with the DevOps team on infrastructure-as-code and observability.

Skills

Apache Kafka
PySpark
Data Lake
DevOps
Python
SQL
AWS

Education

Bachelor's degree in Computer Science/IT or equivalent

Tools

Docker
Kubernetes
CI/CD
Parquet

Job description

  • Design, build and maintain real-time and batch data pipelines for market data ingestion
  • Write distributed processing jobs in PySpark for large-scale transformation and aggregation
  • Ensure pipelines are fault-tolerant, idempotent and recover cleanly from failures
Data lake and storage
  • Design and maintain the data lake, including partitioning strategy, file formats and retention
  • Model data for both analytical querying and low-latency platform reads
  • Manage schema evolution without breaking downstream consumers
Data quality and reliability
  • Build validation and reconciliation checks so bad or missing market data is caught before it reaches subscribers
  • Monitor pipeline health, set up alerting, and own incident resolution and root-cause analysis
  • Maintain lineage and documentation for critical datasets
DevOps and deployment
  • Own deployment of data services using Docker and Kubernetes
  • Build and maintain CI/CD pipelines for data jobs
  • Manage cloud infrastructure (AWS) for data workloads, with an eye on cost
  • Work with the DevOps team on infrastructure-as-code and observability
Data pipelines and streaming
  • Design, build and maintain real-time and batch data pipelines for market data ingestion
  • Build streaming pipelines using Apache Kafka — topics, partitioning, consumer groups, schema management
  • Write distributed processing jobs in PySpark for large-scale transformation and aggregation
  • Ensure pipelines are fault-tolerant, idempotent and recover cleanly from failures
Data lake and storage
  • Design and maintain the data lake, including partitioning strategy, file formats and retention
  • Model data for both analytical querying and low-latency platform reads
  • Manage schema evolution without breaking downstream consumers
Data quality and reliability
  • Build validation and reconciliation checks so bad or missing market data is caught before it reaches subscribers
  • Monitor pipeline health, set up alerting, and own incident resolution and root-cause analysis
  • Maintain lineage and documentation for critical datasets
DevOps and deployment
  • Own deployment of data services using Docker and Kubernetes
  • Build and maintain CI/CD pipelines for data jobs
  • Manage cloud infrastructure (AWS) for data workloads, with an eye on cost
  • Work with the DevOps team on infrastructure-as-code and observability
Requirements
  • 3+ years in data engineering or a closely related role
  • Apache Kafka — production experience with streaming pipelines
  • PySpark — batch and streaming, with an understanding of how to tune a job
  • Data lake architecture — partitioning, formats such as Parquet, and query performance
  • DevOps practices — Docker, Kubernetes, CI/CD
  • Strong Python and SQL
  • Cloud platform experience, preferably AWS (S3, EMR, Glue or equivalent)
  • Bachelor's degree in Computer Science, Information Technology or related, or equivalent practical experience
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Data Enginner
Data Enginner

Zoho • Delhi

On-site
INR 1,200,000 - 2,100,000
Data Engineer
Data Engineer

Coretek Services India • Hyderabad

Hybrid
INR 2,800,000 - 4,000,000
Senior Data Engineer
Senior Data Engineer

Algoleap Technologies • Bengaluru

On-site
INR 1,200,000 - 1,800,000
Data Engineer
Data Engineer

Authorasist Hyderabad • Hyderabad, Bengaluru

Hybrid
INR 1,500,000 - 2,500,000
Data Engineer
Data Engineer

Advance Career Solutions • Pune District, Chennai District, Bengaluru

On-site
INR 1,200,000 - 2,800,000
Data Engineer
Data Engineer

Bacancy Technology Inc • Ahmedabad District

On-site
INR 1,500,000 - 3,000,000
Data Engineer
Data Engineer

fluid.live • Chennai District

On-site
INR 1,200,000 - 1,800,000
Data Engineer
Data Engineer

Coretek Services • Kondapur

On-site
INR 1,800,000 - 3,200,000
Data Engineer
Data Engineer

Scaletrix.AI • Gurugram District

On-site
INR 1,200,000 - 2,400,000
Senior Engineer - Data Engineering
Senior Engineer - Data Engineering

KSB Company • Maharashtra

On-site
INR 600,000 - 1,000,000