Data Engineer Pyspark and Mongo DB

Aligned Automation

Maharashtra

On-site

INR 1,200,000 - 2,400,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Databricks experience
Cloud platforms (Azure/AWS/GCP)

Job summary

Aligned Automation is seeking an experienced Data Engineer with strong MongoDB, PySpark, and Python skills to design, develop and optimize scalable data pipelines. The ideal candidate should handle large datasets, NoSQL databases, distributed data processing and cloud-based data platforms.

You will design, build, and maintain ETL/ELT pipelines, work with MongoDB for data modeling and performance tuning, and optimize Spark jobs for high-performance processing of big data.

Qualifications

  • Python programming
  • PySpark and Spark SQL
  • MongoDB CRUD, indexing, aggregation, replication, sharding
  • SQL skills
  • Git version control
  • Linux/Unix commands
  • JSON, XML, Parquet formats

Responsibilities

  • Design, build, and maintain scalable ETL/ELT pipelines using PySpark and Python.
  • Develop data ingestion frameworks for structured, semi-structured, and unstructured data.
  • Work with MongoDB for data modeling, indexing, and performance tuning.
  • Optimize Spark jobs for high-performance processing of large datasets.
  • Build reusable data transformation and validation frameworks.
  • Develop REST API integrations and automate data ingestion via Python.
  • Monitor and optimize data pipelines for reliability and performance.
  • Collaborate with analysts, data scientists, and application teams to deliver data solutions.
  • Implement data quality, governance, and security best practices.
  • Participate in code reviews and adhere to CI/CD and Agile practices.

Job description

Job Title: Data Engineer (MongoDB, PySpark & Python)

Experience

5-8 Years

Location

As per business requirement

Job Summary

We are looking for an experienced Data Engineer with strong expertise in MongoDB, PySpark, and Python to design, develop, and optimize scalable data pipelines. The ideal candidate should have experience working with large datasets, NoSQL databases, distributed data processing, and cloud-based data platforms.

Key Responsibilities
  • Design, build, and maintain scalable ETL/ELT data pipelines using PySpark and Python.
  • Develop data ingestion frameworks to process structured, semi-structured, and unstructured data.
  • Work extensively with MongoDB for data modeling, querying, indexing, aggregation, and performance optimization.
  • Optimize Spark jobs for high-performance processing of large datasets.
  • Build reusable data transformation and validation frameworks.
  • Develop REST API integrations and automate data ingestion using Python.
  • Monitor, troubleshoot, and optimize data pipelines for reliability and performance.
  • Collaborate with business analysts, data scientists, and application teams to deliver data solutions.
  • Implement data quality, governance, and security best practices.
  • Participate in code reviews and follow CI/CD and Agile development practices.
Required Skills
  • Strong experience in Python programming.
  • Hands-on experience with PySpark and Spark SQL.
  • Strong knowledge of MongoDB, including:
    • CRUD Operations
    • Aggregation Framework
    • Indexing
    • Replication
    • Sharding
    • Performance Tuning
  • Good understanding of data structures and algorithms.
  • Experience in developing ETL/ELT pipelines.
  • Strong SQL skills.
  • Experience with Git version control.
  • Knowledge of Linux/Unix commands.
  • Experience working with JSON, XML, and Parquet data formats.
Preferred Skills
  • Experience with Databricks.
  • Experience with cloud platforms such as Azure, AWS, or GCP.
  • Knowledge of Apache Kafka or other streaming technologies.
  • Experience with orchestration tools such as Apache Airflow.
  • Understanding of Delta Lake and Lakehouse architecture.
  • Familiarity with CI/CD pipelines.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Engineer Pyspark and Mongo DB
Data Engineer Pyspark and Mongo DB

Aligned Automation, LLC • Pune District

On-site
INR 1,500,000 - 2,300,000
Data Engineer
Data Engineer

HGS • Hyderabad

On-site
INR 1,200,000 - 2,100,000
Data Engineer_Spark/Scala
Data Engineer_Spark/Scala

Zorba AI • Kolkata District

On-site
INR 1,000,000 - 1,500,000
Data Engineer - ETL
Data Engineer - ETL

Forward Eye Technologies • Pune District

On-site
INR 2,000,000 - 3,200,000
Data Engineer - SQL/PySpark
Data Engineer - SQL/PySpark

Forward Eye Technologies • Dadri

Hybrid
INR 1,400,000 - 2,000,000
Data Engineer
Data Engineer

Intact Green Services (india) • Bengaluru

On-site
INR 1,800,000 - 2,800,000
Industry-standard compensation
Data Engineer_Spark/Scala
Data Engineer_Spark/Scala

Zorba AI • Maharashtra

On-site
INR 800,000 - 1,200,000
Data Engineer - ETL/PySpark
Data Engineer - ETL/PySpark

Forward Eye Technologies • Dadri

On-site
INR 1,800,000 - 2,400,000
Data Engineer
Data Engineer

Advance Career Solutions • Pune District, Chennai District, Bengaluru

Hybrid
INR 1,200,000 - 2,800,000
Data Engineer_Spark/Scala
Data Engineer_Spark/Scala

Zorba AI • Mumbai

On-site
INR 1,000,000 - 1,500,000