Senior PySpark Data Engineer

Synechron Technologies

Bengaluru

On-site

INR 2,500,000 - 4,200,000

Full time

13 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Synechron Technologies in Bengaluru is seeking a Senior PySpark Data Engineer to design, develop, test, deploy, and support scalable ETL pipelines using Python, PySpark, SQL, and related technologies.

You will build data marts, ensure data quality, and collaborate with multiple teams across the full software development lifecycle, including UAT and production support. 7+ years in data engineering and strong Python/PySpark expertise are required.

Qualifications

  • 7+ years of overall experience in data engineering or related tech roles.
  • 5+ years of commercial experience in a data-driven role.
  • Hands-on experience building data marts and ETL pipelines.
  • Strong Python and PySpark skills for ETL scripting.
  • Experience with Spark, Hadoop, MapReduce, Hive, and Pandas.
  • Strong SQL and Oracle query development; familiarity with NoSQL databases.
  • Experience across the end-to-end SDLC including UAT, production deployment, and post-production support.
  • Ability to process structured, semi-structured, and unstructured data; data quality and validation.

Responsibilities

  • Design, develop, test, deploy, and support scalable ETL pipelines and data marts using Python and PySpark.
  • Build data processing solutions for structured, semi-structured, and unstructured data.
  • Develop clean, maintainable Python and PySpark code.
  • Translate requirements into data engineering solutions and optimize SQL/Oracle queries.
  • Integrate data from multiple sources and ensure data quality and governance.
  • Collaborate with stakeholders and participate in code reviews, testing, and deployment planning.
  • Identify and implement performance, automation, and reliability improvements.

Skills

Data analysis
Communication
Problem solving
Agile / SDLC

Tools

Python
PySpark
Spark
Hadoop
MapReduce
Hive
Pandas
SQL
Oracle
NoSQL
Git
CI/CD

Job description

PySpark Data Engineer – Python, ETL & Data Warehousing Job Summary

Synechron is seeking a PySpark Data Engineer with 7+ years of overall experience and at least 5+ years of commercial experience in data-driven roles . The role will design, develop, test, deploy, and support scalable data pipelines, data marts, and data warehousing solutions using Python, PySpark, SQL, and related data technologies.The position will contribute to business objectives by delivering reliable data solutions, improving data quality and accessibility, supporting analytics and reporting, and ensuring effective data processing across the full software development lifecycle.

Software Requirements
  • Required 7+ years of overall professional experience in data engineering, software development, or related technology roles.
  • Required 5+ years of commercial experience in a data-driven role.
  • Required Hands-on experience building data marts and ETL pipelines.
  • Required Strong expertise in Python and PySpark for ETL scripting.
  • Required Experience writing clean, maintainable, robust, and testable Python code.
  • Required Hands-on experience with Spark, PySpark, Hadoop, MapReduce, Hive, and Pandas.
  • Required Strong knowledge of SQL and Oracle query development.
  • Required Experience working with SQL and NoSQL database management systems.
  • Required Experience across the end-to-end software development lifecycle, including: Build and development.
  • Required Experience across the end-to-end software development lifecycle, including: User acceptance testing.
  • Required Experience across the end-to-end software development lifecycle, including: UAT defect resolution.
  • Required Experience across the end-to-end software development lifecycle, including: Production deployment.
  • Required Experience across the end-to-end software development lifecycle, including: Post-production support.
  • Required Experience debugging PySpark code and investigating data processing issues.
  • Required Strong understanding of data warehousing and data pipeline production practices.
  • Required Ability to process structured, semi-structured, and unstructured data.
  • Required Familiarity with Git, CI/CD processes, data testing, and validation.
  • Required Experience with data analysis, data cleansing, data linking, imputation, and feature engineering.
  • Required Familiarity with workflow orchestration and scheduling tools.
  • Required Experience collaborating with multiple technical and business teams.
  • Preferred Experience with Apache Airflow, Oozie, and Jenkins pipelines.
  • Preferred Experience using Jupyter for data exploration, prototyping, and analysis.
  • Preferred Knowledge of cloud-based data engineering platforms and services.
  • Preferred Experience with data lake, lakehouse, distributed processing, and streaming concepts.
  • Preferred Familiarity with automated data quality monitoring and pipeline observability.
  • Preferred Experience in banking, financial services, or other regulated, data-intensive industries.
  • Preferred Knowledge of data governance, metadata management, lineage, security, and privacy practices.
  • Preferred Experience leading technical workstreams or coordinating delivery across multiple teams.
Overall Responsibilities
  • Design, develop, test, deploy, and support scalable ETL pipelines and data marts using Python and PySpark.
  • Build data processing solutions for structured, semi-structured, and unstructured data.
  • Develop clean, maintainable, robust, and reusable Python and PySpark code.
  • Analyze business and technical requirements and translate them into data engineering solutions.
  • Develop and optimize SQL and Oracle queries for data extraction, transformation, validation, and analysis.
  • Integrate data from multiple sources, databases, files, and systems.
  • Apply data cleansing, data linking, imputation, transformation, validation, and feature engineering techniques.
  • Support data warehouse development, data modeling, data integration, and reporting requirements.
  • Participate in build, UAT, UAT defect resolution, production deployment, and post-production support activities.
  • Debug PySpark code, investigate pipeline failures, and resolve data quality and processing issues.
  • Validate data outputs, reconcile results, and ensure that pipelines meet defined quality and business requirements.
  • Collaborate with technical and non-technical stakeholders to clarify requirements, resolve dependencies, and deliver agreed outcomes.
  • Participate in code reviews, technical discussions, testing, deployment planning, and production support activities.
  • Identify opportunities to improve pipeline performance, automation, reliability, maintainability, and resource efficiency.
  • Maintain technical documentation covering data flows, pipeline logic, data models, dependencies, test evidence, and operational procedures.
  • Consider security, data privacy, cost management, and sustainability when designing and operating data solutions.
Technical Skills (By Category)
Programming Languages
  • Essential Python using a current and supported version.
  • Essential PySpark for distributed data processing and ETL development.
  • Essential Strong understanding of Python functions, modules, object-oriented programming, exception handling, testing, and package management.
  • Essential Ability to write clean, maintainable, robust, reusable, and testable code.
  • Essential SQL for data extraction, transformation, validation, analysis, and query optimization.
  • Essential Understanding of data structures, algorithms, and software engineering principles.
  • Preferred Shell scripting for automation and operational support.
  • Preferred Experience developing reusable Python packages and data-processing utilities.
  • Preferred Knowledge of programming practices for distributed and production-scale data applications.
Databases / Data Management
  • Essential Strong knowledge of relational databases and Oracle query development.
  • Essential Experience with SQL and NoSQL database management systems.
  • Essential Understanding of data warehousing, data marts, data modeling, and data integration.
  • Essential Knowledge of structured, semi-structured, and unstructured data processing.
  • Essential Experience with data cleansing, data linking, imputation, reconciliation, transformation, and validation.
  • Essential Understanding of data quality, data integrity, data lifecycle, and metadata requirements.
  • Essential Ability to analyze large datasets and identify data inconsistencies or processing issues.
  • Preferred Experience with dimensional modeling, fact and dimension tables, and analytical data warehouse design.
  • Preferred Knowledge of data lake and lakehouse architectures.
  • Preferred Familiarity with data lineage, metadata management, and data governance.
  • Preferred Experience with feature engineering and preparing data for analytics or machine learning use cases.
  • Preferred Knowledge of database performance tuning and query optimization.
Cloud Technologies
  • Essential Understanding of cloud-based data engineering concepts and distributed data processing.
  • Essential Awareness of cloud storage, compute, networking, access management, monitoring, and deployment considerations.
  • Essential Ability to support data pipelines across development, test, UAT, and production environments.
  • Preferred Experience developing and deploying PySpark data pipelines on cloud platforms.
  • Preferred Familiarity with cloud-based data lakes, data warehouses, managed databases, and workflow services.
  • Preferred Knowledge of cloud monitoring, infrastructure automation, identity management, and security controls.
  • Preferred Understanding of cost-efficient and sustainable use of cloud data-processing resources.
Frameworks and Libraries
  • Essential Apache Spark and PySpark.
  • Essential Hadoop, MapReduce, and Hive.
  • Essential Pandas for data analysis and transformation.
  • Essential Python libraries for database connectivity, file handling, data validation, and automation.
  • Essential Experience developing ETL and data-processing frameworks.
  • Essential Understanding of distributed processing, partitioning, transformations, actions, and performance considerations.
  • Preferred Apache Airflow or Oozie for workflow orchestration.
  • Preferred Jupyter for data analysis, exploration, and prototyping.
  • Preferred Libraries supporting data quality, testing, feature engineering, and statistical analysis.
  • Preferred Familiarity with streaming or near-real-time data-processing frameworks.
Development Tools and Methodologies
  • Essential Experience across the end-to-end SDLC, including build, UAT, defect fixing, deployment, and post-production support.
  • Essential Git for source code versioning, branching, merging, and code review.
  • Essential Familiarity with CI/CD processes and automated build or deployment workflows.
  • Essential Experience with data testing, validation, reconciliation, and defect management.
  • Essential Knowledge of Agile or iterative software delivery practices.
  • Essential Ability to document data flows, transformation logic, data dependencies, test results, and operational procedures.
  • Essential Experience coordinating with multiple teams to resolve dependencies and deliver project outcomes.
  • Preferred Jenkins pipeline experience.
  • Preferred Experience with automated data quality checks and test execution.
  • Preferred Familiarity with pipeline monitoring, logging, alerting, and incident management.
  • Preferred Knowledge of infrastructure as code and automated environment deployment.
  • Preferred Experience with performance monitoring and optimization of production data pipelines.
Security Protocols
  • Essential Understanding of secure data handling and data protection principles.
  • Essential Awareness of authentication, authorization, identity and access management, encryption, secrets management, and secure connectivity.
  • Essential Ability to apply appropriate access controls to data pipelines, databases, files, and processing environments.
  • Essential Understanding of data privacy, data integrity, auditability, and secure transfer practices.
  • Preferred Experience implementing security controls across cloud and on-premises data environments.
  • Preferred Knowledge of data masking, tokenization, role-based access control, and audit logging.
  • Preferred Familiarity with vulnerability management, security testing, and compliance-related data controls.
  • Preferred Understanding of secure configuration and monitoring practices for data platforms.
Experience Requirements
  • 7+ years of overall experience in data engineering, software development, or related technology roles.
  • 5+ years of commercial experience in a data-driven role.
  • Experience building data marts and ETL pipelines.
  • Strong hands-on experience with Python and PySpark for ETL scripting.
  • Experience with Spark, Hadoop, MapReduce, Hive, Pandas, SQL, and Oracle queries.
  • Experience working with SQL and NoSQL database technologies.
  • Experience across build, UAT, UAT defect resolution, production deployment, and post-production support.
  • Experience debugging PySpark code and resolving data pipeline, data quality, and production issues.
  • Strong understanding of data warehousing and production data pipeline practices.
  • Experience handling structured, semi-structured, and unstructured data.
  • Experience with CI/CD, Git, data testing, validation, workflow scheduling, and pipeline support.
  • Experience in banking, financial services, or other regulated data-intensive industries is preferred.
  • Candidates may also qualify through equivalent practical experience, relevant certifications, professional training, or demonstrated delivery of complex data engineering solutions.
Day-to-Day Activities
  • Design, develop, test, and maintain Python and PySpark ETL pipelines, data marts, data transformations, and data warehouse components.
  • Collaborate with technical and non-technical stakeholders, participate in Agile meetings, clarify requirements, and resolve cross-team dependencies.
  • Perform data analysis, Oracle query development, PySpark debugging, data validation, UAT defect fixing, and production support.
  • Review pipeline results, monitor delivery progress, document technical outcomes, recommend improvements, and make implementation decisions within approved standards.
Qualifications
  • Degree in Computer Science, Information Technology Engineering, or an equivalent discipline; equivalent professional experience may be considered.
  • Minimum of 7+ years of overall experience , including at least 5+ years of commercial experience in data-driven roles.
  • Certifications in data engineering, cloud technologies, Python, Spark, Agile, or database technologies are preferred.
  • Complete Synechron-required training related to information security, data protection, data governance, workplace conduct, and responsible technology use.
  • Maintain continuous professional development in Python, PySpark, Spark, data warehousing, cloud data platforms, SQL, automation, security, and data engineering practices.
Professional Competencies
  • Critical thinking, data analysis, technical investigation, and structured problem-solving.
  • Technical ownership, teamwork, dependency coordination, and delivery accountability.
  • Clear communication with technical and non-technical stakeholders.
  • Adaptability, continuous learning, and effective response to changing data and delivery requirements.
  • Innovation focused on reliable, maintainable, automated, scalable, and sustainable data solutions.
  • Effective prioritization, organization, time management, and delivery under multiple deadlines.
S YNECHRON’S DIVERSITY & INCLUSION STATEMENT

Diversity & Inclusion are fundamental to our culture, and Synechron is proud to be an equal opportunity workplace and is an affirmative action employer. Our Diversity, Equity, and Inclusion (DEI) initiative ‘Same Difference’ is committed to fostering an inclusive culture – promoting equality, diversity and an environment that is respectful to all. We strongly believe that a diverse workforce helps build stronger, successful businesses as a global company. We encourage applicants from across diverse backgrounds, race, ethnicities, religion, age, marital status, gender, sexual orientations, or disabilities to apply. We empower our global workforce by offering flexible workplace arrangements, mentoring, internal mobility, learning and development programs, and more. All employment decisions at Synechron are based on business needs, job requirements and individual qualifications, without regard to the applicant’s gender, gender identity, sexual orientation, race, ethnicity, disabled or veteran status, or any other characteristic protected by law .

Candidate Application Notice Experience Level Senior Level

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

PySpark Data Engineer – Python, ETL & Data Warehousing
PySpark Data Engineer – Python, ETL & Data Warehousing

Synechron • India

On-site
INR 2,000,000 - 3,500,000
Senior Data Engineer - AWS, Spark, Scala, Big Data & GenAI
Senior Data Engineer - AWS, Spark, Scala, Big Data & GenAI

Synechron Technologies Pvt. Ltd._INDIA Company • Bengaluru

On-site
INR 2,500,000 - 4,000,000
Senior Data Engineer - AWS, Spark, Scala, Big Data & GenAI
Senior Data Engineer - AWS, Spark, Scala, Big Data & GenAI

Synechron • India

On-site
INR 1,500,000 - 3,000,000
Python Full Stack Developer
Python Full Stack Developer

Synechron Technologies • Bengaluru

On-site
INR 1,400,000 - 2,000,000
DevOps Engineer - Big Data & Hadoop with Kubernetes and CI/CD
DevOps Engineer - Big Data & Hadoop with Kubernetes and CI/CD

Synechron Technologies • Bengaluru

On-site
INR 2,800,000 - 4,500,000
Java, React & Python Full Stack Developer – Cloud, APIs & Agile Delivery
Java, React & Python Full Stack Developer – Cloud, APIs & Agile Delivery

Synechron • Bengaluru

On-site
INR 3,500,000 - 6,000,000
Flexible workplace arrangements
Mentoring and development programs
Python AI Developer – Machine Learning, Cloud & Agile Delivery
Python AI Developer – Machine Learning, Cloud & Agile Delivery

Synechron • Bengaluru

On-site
INR 3,000,000 - 6,000,000
Full Stack Developer (Java, React & Python) - Cloud and Agile Delivery
Full Stack Developer (Java, React & Python) - Cloud and Agile Delivery

Synechron Technologies • Bengaluru

On-site
INR 3,200,000 - 4,000,000
Java Developer – Cloud, Microservices & Agile Delivery
Java Developer – Cloud, Microservices & Agile Delivery

Synechron • Hyderabad

On-site
INR 3,600,000 - 5,400,000
PySpark Developer
PySpark Developer

Synechron Inc. • Bengaluru

On-site
INR 1,200,000 - 1,800,000
Flexible workplace arrangements
Learning and development programs