Senior Data Engineer - Pyspark & AWS

RBM Software

Pune District

On-site

INR 4,000,000 - 7,000,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

RBM Software is seeking an experienced Senior Data Engineer to own and scale our data platform in Pune, India. You will drive PySpark processing, AWS-based pipelines, and reliable data lake architecture, balancing development with production support.

You will mentor junior teammates and evolve our observability and incident response practices. You will orchestrate Airflow DAGs, build serverless ingestion with Glue and Lambda, and structure secure S3 data lakes with clear landing, raw, and gold

Qualifications

  • 5+ years in data engineering or platform support with PySpark
  • Strong AWS ecosystem experience (3+ years)
  • Bachelor’s or Master’s in CS/SE or equivalent
  • Analytical communicator capable of translating roadblocks into technical fixes

Responsibilities

  • Orchestrate complex DAGs with Apache Airflow
  • Build serverless data pipelines using AWS Glue and PySpark
  • Deploy cost-optimized AWS Lambda scripts for event-driven tasks
  • Manage S3 data lakes with landing, raw, and gold tiers
  • Provide end-to-end platform ownership, monitoring, and incident handling
  • Mentor junior engineers through code reviews and guidance
  • Establish engineering docs, tests, and on-call rotation standards

Skills

AWS Production Environment
Observability & Triages
Serverless Ingestion
Data Lakes & Orchestration
Apache Spark
Python Programming
Database Foundations

Education

Bachelor’s or Master’s degree in Computer Science, Software Engineering, or equivalent

Tools

Apache Airflow
AWS Glue
AWS Lambda
Amazon S3
Amazon Redshift
Amazon Athena
Amazon RDS
AWS CloudWatch
CloudTrail
X-Ray

Job description

Work Mode: Work from Office (5 days working)

Job Summary

We are seeking an experienced and leadership-minded Senior Data Engineer - Pyspark & AWS to drive the reliability, scaling, and operation of our enterprise data infrastructure. In this role, you will balance core development with production support and operational excellence, ensuring high availability through proactive monitoring and debugging of cloud data streams.

As a senior member of the data team, you will take full ownership of the data ecosystem, troubleshoot complex live incidents, and actively mentor and upskill junior engineers.

Key Responsibilities
  • Orchestration: Program complex workflow dependencies and multi-stage DAGs within Apache Airflow.
  • ETL Architecture: Construct serverless data preparation workflows using AWS Glue and PySpark.
  • Compute Processing: Deploy targeted, cost-optimized scripts via AWS Lambda for event-driven backend tasks.
  • Storage Management: Structure secure Amazon S3 data lakes with clear landing, raw, and gold curation tiers.
  • Technical Ownership: Accept end-to-end accountability for platform health, security compliance, and architecture updates.
  • Team Mentorship: Guide junior developers through systematic code reviews, architectural advice, and pair programming sessions.
  • Process Standards: Establish baseline engineering documentation, testing frameworks, and on-call rotation protocols.
Technical Skills
Primary Skills (Required)
  • AWS Production Environment: Background supporting enterprise-scale cloud infrastructures under tight data freshness targets.
  • Observability & Triage: Hands-on triage using AWS CloudWatch, CloudTrail, X-Ray, or third-party log collectors.
  • Serverless Ingestion: Advanced setup of AWS Glue (Crawlers, Catalogs, Jobs) and event-driven AWS Lambda components.
  • Data Lakes & Orchestration: Deep mastery of Amazon S3 object storage policies and workflow automation with Apache Airflow.
  • Apache Spark: Processing large-scale datasets using distributed computing engines via PySpark or Spark SQL.
  • Python Programming: Advanced scripting for data automation, API bindings, custom transformations, and object mutation.
  • Database Foundations: Ability to query and locate performance degradation points inside Amazon Redshift, Athena, or RDS systems.
Production Support & Operations
  • Production Support: Provide Tier 3 technical support for live cloud datasets, maintaining strict platform uptime SLAs.
  • System Monitoring: Build dashboards using AWS CloudWatch to track resource consumption and data flow metrics.
  • Issue Debugging: Inspect execution logs, trace data failures, and resolve runtime bottlenecks across active microservices.
  • Incident Management: Lead root-cause analysis (RCA) activities for data dropouts, pipeline breaks, or structural data quality shifts.
Qualifications & Experience
  • Experience: 5+ years of dedicated data engineering or platform support experience, with strong Pyspark programming skills and at least 3 years directly focused on AWS ecosystems.
  • Education: Bachelor’s or Master’s degree in Computer Science, Software Engineering, or an equivalent technical discipline.
  • Soft Skills: Highly analytical communicator capable of translating business roadblocks into immediate technical fixes.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Data Engineer
Senior Data Engineer

AagatiServe Pvt Ltd • Delhi

On-site
INR 1,800,000 - 2,400,000
AWS + Pyspark Data Engineer ( Gurugram)
AWS + Pyspark Data Engineer ( Gurugram)

PwC India • Gurugram District

Hybrid
INR 1,400,000 - 2,100,000
Senior Data Engineer
Senior Data Engineer

Arcana Analytics • Bengaluru

Hybrid
INR 1,200,000 - 1,800,000
Senior Data Engineer (AWS)
Senior Data Engineer (AWS)

Objectways • Bengaluru

On-site
INR 1,800,000 - 2,500,000
Data Engineer
Data Engineer

Calsoft • Pune District

On-site
INR 2,800,000 - 4,500,000
Senior Data Engineer
Senior Data Engineer

Masscom Corporation • Gujarat

On-site
INR 4,000,000 - 7,000,000
Senior Data Engineer
Senior Data Engineer

SourcingXPress • Hyderabad

On-site
INR 1,500,000 - 2,500,000
Data Engineer Pyspark and Mongo DB
Data Engineer Pyspark and Mongo DB

Aligned Automation • Maharashtra

On-site
INR 1,200,000 - 2,400,000
Databricks experience
Cloud platforms (Azure/AWS/GCP)
Lead AWS Data Engineer
Lead AWS Data Engineer

Weekday (YC W21) • Bengaluru

On-site
INR 1,000,000 - 1,500,000
Senior Data Engineer
Senior Data Engineer

Fulcrum Worldwide Software • Pune District

Hybrid
INR 1,200,000 - 2,400,000