Pyspark Developer ( Mumbai)

V2 Solutions

Hinoba-an

On-site

PHP 796,000 - 1,194,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

V2 Solutions is seeking a Big Data Technical Senior Associate Engineer to design and deliver scalable data products on Cloudera platforms. You will partner with Business, Risk, and Technology teams in the BFSI domain, building robust pipelines and ensuring data quality in batch and streaming modes.

The role emphasizes hands-on expertise with Hadoop/Spark, SQL, and cloud/enterprise data platforms, with on-site Mumbai as the location. Prior experience in onsite/offshore models is valued.

Qualifications

  • Must have hands-on experience building data pipelines on Hadoop/Spark ecosystem.
  • Strong Spark (Scala and/or PySpark), Hive/Impala SQL, and performance tuning.
  • Kafka for streaming ingestion and NiFi/StreamSets for batch/near-real-time flows.
  • Experience with Cloudera Manager, YARN/Tez, HDFS, and Oozie/Airflow.
  • Data warehousing concepts: dimensional modelling, partitioning, bucketing.
  • Linux/Unix, Shell scripting, Git, and CI/CD (Jenkins/GitLab CI).
  • Strong SQL and data modelling for BFSI use cases (lending, liabilities, risk, reporting).
  • Experience in writing HLD/LLD and unit/integration testing.
  • Exposure to SDLC/Agile and onsite/offshore model.

Responsibilities

  • Develop robust Spark jobs (batch and streaming) with unit tests and observability.
  • Implement ingestion patterns (Kafka/NiFi), data quality checks, and job scheduling.
  • Analyze and tune SQL/Spark for large-scale datasets.

Skills

Hadoop
Spark
Hive/Impala SQL
Kafka
NiFi/StreamSets
Cloudera Manager/YARN/Tez
HDFS
Oozie/Airflow
Data warehousing modeling
Linux/Unix
Shell scripting
Git
CI/CD (Jenkins/GitLab CI)
SQL
BFSI data modelling
HLD/LLD docs
Agile

Education

B.E./B.Tech or equivalent
Cloudera Certified Associate/Professional or CDP certifications
AWS/Azure/GCP data certifications (nice to have)
Databricks Lakehouse Fundamentals or Associate (nice to have)

Tools

Cloudera Data Platform
Docker
Kubernetes

Job description

Big Data Technical Senior Associate Engineer

We are looking for a suitable candidate for the opening of the Senior/Associate Engineer role with good to have experience in the Banking and Financial Services domain with 4-6 years of relevant experience in Data Engineering / Big Data platforms. The candidate will work closely with Business, Risk, and Technology stakeholders to deliver scalable data products on Cloudera (CDH/CDP).

Key Roles & Responsibilities
  • Must have / Primary Skills / Mandatory Hands-on experience building data pipelines on Hadoop/Spark ecosystem
  • Strong Spark (Scala and/or PySpark), Hive/Impala SQL, and performance tuning
  • Working knowledge of Kafka for streaming ingestion and NiFi (or StreamSets) for batch/near-real-time flows
  • Experience with Cloudera Manager, YARN/Tez, HDFS, and job orchestration using Oozie/Airflow
  • Good understanding of data warehousing concepts (dimensional modelling, partitioning, bucketing)
  • Proficiency in Linux/Unix, Shell scripting, Git, and CI/CD (Jenkins/GitLab CI)
  • Strong SQL and data modelling for BFSI use cases (lending, liabilities, risk, regulatory reporting)
  • Experience in writing technical design documents (HLD/LLD) and unit/integration testing
  • Exposure to SDLC/Agile and working in onsiteoffshore model

Location - Mumbai

  • Develop robust Spark jobs (batch and streaming) with unit tests and observability
  • Implement ingestion patterns (Kafka/NiFi), data quality checks, and job scheduling
  • Analyze and tune SQL/Spark for large-scale datasets
Good to have / Secondary Skills / Desired Experience
  • with Cloudera Data Platform (CDP)
  • Private Cloud Base/Public Cloud Security and governance: Apache Ranger, Atlas; Kerberos; Sentry (legacy)
  • Cloud data services – AWS (EMR, Glue, S3), Azure (HDInsight, Synapse, ADLS), or GCP (Dataproc, BigQuery)
  • Databricks experience (Spark, Delta Lake) for select workloads
  • Containerization and orchestration (Docker/Kubernetes) for micro-batch/ML workloads
  • Python for data processing and utilities; familiarity with Scala build tools (sbt/maven)
  • Monitoring/observability – Cloudera Manager metrics, Grafana/Prometheus, log aggregation
  • Experience with BI consumption patterns and semantic layers for risk/regulatory dashboards
Educational Qualifications
  • B.E./B.Tech or equivalent
Experience Range 4-6 years
Certifications (Preferred)
  • Cloudera Certified Associate/Professional (CCA/CCP) or CDP certifications
  • AWS/Azure/GCP data certifications (nice to have)
  • Databricks Lakehouse Fundamentals or Associate (nice to have)
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior PySpark Data Engineer – Big Data (BFSI)
Senior PySpark Data Engineer – Big Data (BFSI)

V2 Solutions • Hinoba-an

On-site
PHP 796,000 - 1,194,000
Software Development Engineer
Software Development Engineer

V2 Solutions • Hinoba-an

On-site
PHP 900,000 - 1,300,000
PySpark Developer
PySpark Developer

V2 Solutions • Hinoba-an

On-site
PHP 600,000 - 900,000
Senior Data Architect
Senior Data Architect

V2 Solutions • Hinoba-an

On-site
PHP 1,200,000 - 1,800,000
Data Platform Engineer
Data Platform Engineer

V2 Solutions • Hinoba-an

On-site
PHP 300,000 - 540,000
Lead Software Engineer – Data
Lead Software Engineer – Data

V2 Solutions • Hinoba-an

On-site
PHP 2,009,000 - 3,125,000
Big Data - Senior Engineer
Big Data - Senior Engineer

Iris Software, Inc. • Hinoba-an

On-site
PHP 1,990,000 - 3,316,000
Databricks Unified Data Analytics Platform Engineer
Databricks Unified Data Analytics Platform Engineer

Accenture in the Philippines • Philippines

On-site
PHP 600,000 - 900,000
Senior Data Engineer
Senior Data Engineer

Bounteous • Hinoba-an

On-site
PHP 800,000 - 1,200,000
DGS India - Bengaluru - Manyata N1 Block / Technology
DGS India - Bengaluru - Manyata N1 Block / Technology

Merkle Schweiz • Hinoba-an

On-site
PHP 900,000 - 1,500,000