Data Engineer Pyspark and Mongo DB

Aligned Automation, LLC

Pune District

On-site

INR 1,500,000 - 2,300,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Aligned Automation Services Pvt Ltd is seeking an experienced Data Engineer to design, develop, and optimize scalable data pipelines. The candidate will leverage MongoDB, PySpark, and Python to handle large datasets and NoSQL data platforms.

Key responsibilities include building ETL/ELT pipelines, data ingestion frameworks, and data quality measures. Collaboration with analysts and data scientists is essential, with CI/CD and Agile practices involved.

Qualifications

  • 5–8 years of data engineering experience.
  • Strong Python and PySpark programming skills.
  • Extensive MongoDB experience: CRUD, indexing, aggregation, replication, sharding, performance tuning.
  • Experience with large-scale data pipelines and ETL/ELT processes.
  • Familiarity with JSON, XML, Parquet data formats and data modeling.

Responsibilities

  • Design, build, and maintain scalable ETL/ELT pipelines using PySpark and Python.
  • Develop data ingestion frameworks for structured, semi-structured and unstructured data.
  • Work with MongoDB for data modeling, querying, and performance optimization.
  • Optimize Spark jobs for high-volume data processing.
  • Create reusable data transformation and validation components.
  • Develop REST API integrations and automate data ingestion using Python.
  • Monitor and troubleshoot pipelines for reliability and performance.
  • Collaborate with analysts, data scientists, and apps teams.

Skills

Python
PySpark
MongoDB
Spark SQL
Git
Linux/Unix
JSON/XML/Parquet

Tools

Databricks
Azure
AWS
GCP
Apache Kafka
Apache Airflow
Delta Lake

Job description

Aligned Automation Services Pvt Ltd | Full time

Experience

5–8 Years

Location

As per business requirement

Job Summary

We are looking for an experienced Data Engineer with strong expertise in MongoDB, PySpark, and Python to design, develop, and optimize scalable data pipelines. The ideal candidate should have experience working with large datasets, NoSQL databases, distributed data processing, and cloud-based data platforms.

Key Responsibilities
  • Design, build, and maintain scalable ETL/ELT data pipelines using PySpark and Python .
  • Develop data ingestion frameworks to process structured, semi-structured, and unstructured data.
  • Work extensively with MongoDB for data modeling, querying, indexing, aggregation, and performance optimization.
  • Optimize Spark jobs for high-performance processing of large datasets.
  • Build reusable data transformation and validation frameworks.
  • Develop REST API integrations and automate data ingestion using Python.
  • Monitor, troubleshoot, and optimize data pipelines for reliability and performance.
  • Collaborate with business analysts, data scientists, and application teams to deliver data solutions.
  • Implement data quality, governance, and security best practices.
  • Participate in code reviews and follow CI/CD and Agile development practices.
Required Skills
  • Strong experience in Python programming.
  • Hands-on experience with PySpark and Spark SQL.
  • Strong knowledge of MongoDB , including:
    • CRUD Operations
    • Aggregation Framework
    • Indexing
    • Replication
    • Sharding
    • Performance Tuning
  • Good understanding of data structures and algorithms.
  • Experience with Git version control.
  • Knowledge of Linux/Unix commands.
  • Experience working with JSON, XML, and Parquet data formats.
Preferred Skills
  • Experience with Databricks .
  • Experience with cloud platforms such as Azure , AWS , or GCP .
  • Knowledge of Apache Kafka or other streaming technologies.
  • Experience with orchestration tools such as Apache Airflow .
  • Understanding of Delta Lake and Lakehouse architecture.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Engineer Pyspark and Mongo DB
Data Engineer Pyspark and Mongo DB

Aligned Automation • Maharashtra

On-site
INR 1,200,000 - 2,400,000
Databricks experience
Cloud platforms (Azure/AWS/GCP)
Pyspark developer
Pyspark developer

Aligned Automation • Pune District

On-site
INR 2,200,000 - 3,400,000
Data Engineer
Data Engineer

Intact Green Services (india) • Bengaluru

On-site
INR 1,800,000 - 2,800,000
Industry-standard compensation
Data Engineer
Data Engineer

Aligned Automation • Maharashtra

On-site
INR 800,000 - 1,200,000
Data Engineer
Data Engineer

EXL • Pune District

On-site
INR 1,200,000 - 2,400,000
Data Engineer
Data Engineer

Keka Technologies Private Limited • Indore District

On-site
INR 800,000 - 1,300,000
Sr. Data Engineer (Databricks)
Sr. Data Engineer (Databricks)

Blumetra Solutions • Hyderabad

Hybrid
INR 1,200,000 - 1,500,000
Data Engineer- Databricks
Data Engineer- Databricks

r3 Consultant • Bengaluru

On-site
INR 1,000,000 - 2,000,000
Data Engineer_Spark/Scala
Data Engineer_Spark/Scala

Zorba AI • Kolkata District

On-site
INR 1,000,000 - 1,500,000
Pyspark Data Engineer
Pyspark Data Engineer

Synechron • Bengaluru

On-site
INR 1,200,000 - 1,800,000