Lead Pyspark Cloud Data Engineer

enGen Global

Hyderabad

Hybrid

INR 3,000,000 - 7,000,000

Full time

2 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

enGen Global is seeking a Lead PySpark Cloud Data Engineer in a hybrid role based in Chennai / Hyderabad. The candidate will design and implement scalable data pipelines using PySpark and Dataproc, building lakehouse architectures with Iceberg and BigLake Metastore.

You will mentor junior engineers and produce robust data models and documentation. Requirements include 8-12 years of experience, strong PySpark, SQL, and cloud experience with GCP services.

Qualifications

  • 8-12 years total experience building high-performant data pipelines (batch and streaming).
  • Strong hands-on PySpark for data processing, transformation and analysis.
  • Experience with cloud data platforms, preferably GCP (Dataproc, BigQuery, Cloud Storage).
  • Knowledge of Apache Iceberg for open table formats.
  • Advanced SQL skills for querying, manipulation and optimization.
  • Proficiency with Git and collaborative development workflows.
  • Excellent analytical and problem-solving abilities with attention to detail.

Responsibilities

  • Design, develop, and implement scalable data pipelines using PySpark, Dataproc, and Google cloud-native technologies.
  • Create pipelines for data lakehouse architectures utilizing Apache Iceberg and BigLake Metastore.
  • Develop and maintain data transformation logic using PySpark and/or dbt.
  • Work with Google Cloud Platform services including Dataproc, BigQuery, Cloud Storage.
  • Optimize data pipelines and queries for performance, efficiency, and cost.
  • Implement data quality checks, monitoring, and governance best practices.
  • Provide technical guidance and mentorship to junior data engineers.
  • Create and maintain comprehensive technical documentation for data pipelines and architecture.
  • Provide regular updates on tasks, status and risks to project manager.

Skills

PySpark
SQL
Git
Data pipelines
BigQuery
Dataproc
dbt
Iceberg
BigLake Metastore

Education

Bachelor’s degree or higher

Tools

Dataproc
Iceberg
BigLake Metastore
dbt
Kafka/Pub-Sub

Job description

Job Title: Lead PySpark Cloud Data Engineer

Work Location: Chennai / Hyderabad

Experience: 8 -12 years

Work Model: Hybrid (3 days WHO)

Shift Time: 1PM to 10PM / 3PM to 12PM

enGen Global is an emerging global healthcare partner that delivers strategic innovation, expertise, and flexibility to its healthcare partners. Being a US healthcare conglomerate captive, we have direct access to deeper insights that help us accelerate our learning process and keeps us ahead of the curve. Thryve delivers next-generation solutions that enable our healthcare partners to provide positive experiences to their consumers.

Our global collaborative of healthcare, operations, and IT experts creates innovative and sustainable processes for our clients, which keeps the ever-evolving consumers engaged and assists them in managing the future of their healthcare better. We recognize that our people are our strength and the diverse talents they bring to our global workforce are directly linked to our success. Thryve is an equal opportunity employer and places a high value on integrity, diversity, and inclusion in the organization. We do not discriminate based on any protected attribute. For more information about the organization, please visit www.thryvedigital.com

Role Summary:

This job involves understand the overall requirement of the enterprise Data need and design & develop the robust, highly scalable & resilient data pipeline using Pyspark, Dataproc, open source table formats and other Google services. Job will involve extensive interfacing & co-ordination with other senior tech folks (across India & US) and lead the design & development of the data ingestion pipelines

  • Design, develop, and implement highly scalable, reliable, and performant data pipelines using PySpark, Dataproc, and other Google cloud-native technologies
  • Create pipelines for data lakehouse architectures utilizing Apache Iceberg for open table formats and BigLake Metastore for unified metadata management
  • Develop and maintain data transformation logic using Pyspark and / or dbt to create clean, consistent, and production-ready data models
  • Work extensively with Google Cloud Platform (GCP) services, including Dataproc, BigQuery, Cloud Storage, and other relevant data services.
  • Optimize data pipelines and queries for performance, efficiency, and cost-effectiveness.
  • Implement and enforce data quality checks, monitoring, and governance best practices to ensure data integrity and reliability
  • Provide technical guidance and mentorship to junior data engineers, fostering a culture of continuous learning and improvement
  • Create and maintain comprehensive technical documentation for data pipelines, data models, and platform architecture
  • Provide regular updates on the tasks, status and risks to project manager
The experience we are looking to add to our team
Required
  • Bachelor’s degree or higher from a reputed university
  • 8 to 12 years total experience with majority of that experience related to building high performant data pipelines – batch and streaming
  • Strong Hands on experience / expert-level proficiency in PySpark for data processing, transformation, and analysis
  • Extensive experience in implementing large scale data ingestion and curation solutions
  • Hands-on experience with cloud-based data platforms, preferably Google Cloud Platform (GCP) and services like Dataproc, BigQuery, and Cloud Storage
  • Knowledge in Apache Iceberg for open table formats
  • Advanced SQL skills for data querying, manipulation, and optimization
  • Proficiency with Git and collaborative development workflows
  • Excellent analytical and problem-solving skills with a keen attention to detail
  • Strong communication and interpersonal skills, with the ability to explain complex technical concepts to both technical and non-technical audiences
Good to have
  • Expertise in Google Cloud services
  • Proven experience with Apache Iceberg for open table formats and BigLake Metastore for unified metadata management
  • Strong experience with dbt (data build tool) for data modeling, transformation, and orchestration
  • Experience in Data Governance and Data quality tools like Atlan, Monte Carlo etc.
  • Experience in processing streaming data using Kafka / Pub-Sub
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Cloud Data Architect
Cloud Data Architect

enGen Global • Hyderabad

On-site
INR 2,500,000 - 4,500,000
Gcp Data Engineer
Gcp Data Engineer

Lloyds Technology Centre • Hyderabad

Hybrid
INR 2,500,000 - 4,200,000
ETL Technical Product Owner / Manager
ETL Technical Product Owner / Manager

enGen Global • Chennai District

On-site
INR 1,500,000 - 2,100,000
Technical Product Manager
Technical Product Manager

enGen Global • Chennai District

On-site
INR 3,500,000 - 7,000,000
PySpark Developer / Senior Data Engineer
PySpark Developer / Senior Data Engineer

Alignity Solutions • Hyderabad

On-site
INR 1,200,000 - 1,800,000
Cloud Data Engineer
Cloud Data Engineer

enGen Global • Chennai District

On-site
INR 1,800,000 - 3,200,000
Data Engineer
Data Engineer

Intact Green Services (india) • Bengaluru

On-site
INR 1,800,000 - 2,800,000
Industry-standard compensation
Consultant - Cloud Data Engineer
Consultant - Cloud Data Engineer

enGen Global • Chennai District

On-site
INR 1,800,000 - 2,400,000
DATA ENGINEER
DATA ENGINEER

Covaicareers.com • Chennai District

Hybrid
INR 2,500,000 - 4,000,000
Data Engineer Pyspark GCP - Devops - Immediate or not working
Data Engineer Pyspark GCP - Devops - Immediate or not working

Kairos Technologies • Kolkata District, Chennai District, Bengaluru

On-site
INR 1,800,000 - 2,800,000