Python, Pyspark Developer

Infosys

Hyderabad

On-site

INR 1,100,000 - 2,300,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Infosys in Hyderabad is seeking a seasoned Data Engineer to design, build, and optimize scalable data pipelines using Python and PySpark. You will implement efficient transformations, write complex SQL, and ensure data quality across multiple sources in distributed environments.

The role requires 5-9 years of hands-on experience in data engineering or software development, with strong Python and PySpark skills, and a degree in CS or Engineering.

Qualifications

  • 5-9 years of hands-on experience in software development and/or data engineering.
  • Strong proficiency in Python with production-grade applications or data workflows.
  • Strong PySpark proficiency, including DataFrame APIs, optimization, and distributed processing.
  • Working knowledge of SQL for complex queries, data analysis, and validation.
  • Bachelor’s degree in CS/Engineering or equivalent practical experience.

Responsibilities

  • Design, develop, and maintain scalable batch/stream data pipelines using Python and PySpark in distributed environments.
  • Implement efficient transformations, aggregations, and joins on large datasets with performance and cost optimization.
  • Write optimized SQL for data extraction, validation, and reconciliation across multiple sources.
  • Build reusable, testable modules and follow engineering best practices (code reviews, unit testing, documentation).
  • Troubleshoot production issues, perform root-cause analysis, and implement long-term fixes and monitoring improvements.
  • Collaborate with stakeholders to translate requirements into technical designs, delivery plans, and measurable outcomes.
  • Ensure data quality through validation checks, anomaly detection patterns, and consistent schema management.
  • Contribute to continuous improvement of development standards, performance benchmarks, and pipeline reliability.

Skills

Python
PySpark
SQL
Data engineering

Education

Bachelor’s degree in Computer Science or Engineering

Job description

Technology->Analytics - Packages->Python - Big Data,Technology->Big Data - Data Processing->PySpark

  • Design, develop, and maintain scalable batch/stream data pipelines using Python and PySpark in distributed environments.
  • Implement efficient transformations, aggregations, and joins on large datasets while ensuring performance and cost optimization.
  • Write optimized SQL for data extraction, validation, and reconciliation across multiple sources.
  • Build reusable, testable modules and follow engineering best practices (code reviews, unit testing, documentation).
  • Troubleshoot production issues, perform root-cause analysis, and implement long-term fixes and monitoring improvements.
  • Collaborate with stakeholders to translate requirements into technical designs, delivery plans, and measurable outcomes.
  • Ensure data quality through validation checks, anomaly detection patterns, and consistent schema management.
  • Contribute to continuous improvement of development standards, performance benchmarks, and pipeline reliability.
  • Bachelor’s degree in Computer Science, Engineering, or a related field (or equivalent practical experience).
  • 5-9 years of hands‑on experience in software development and/or data engineering roles.
  • Strong proficiency in Python with experience building production‑grade applications or data workflows.
  • Strong proficiency in PySpark, including DataFrame APIs, optimization techniques, and distributed processing concepts.
  • Working knowledge of SQL for complex queries, data analysis, and validation.
  • Experience delivering reliable solutions with attention to performance, scalability, and maintainability.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Developer - PySpark
Developer - PySpark

Compunnel, Inc. • Pune District

On-site
INR 800,000 - 1,500,000
Python-Pyspark Developer
Python-Pyspark Developer

Infosys • Hyderabad

On-site
INR 2,400,000 - 3,600,000
Python, PySpark, ETL Developer
Python, PySpark, ETL Developer

Infosys • Hyderabad

On-site
INR 800,000 - 1,600,000
Developer
Developer

GSB Solutions • Hyderabad

On-site
INR 1,200,000 - 1,500,000
Data Engineer (Python & PySpark)
Data Engineer (Python & PySpark)

Techknomatic Services • Pune District

On-site
INR 700,000 - 1,200,000
Pyspark Data Engineer
Pyspark Data Engineer

Synechron • Bengaluru

On-site
INR 1,200,000 - 1,800,000
Python, Spark Scala Developer
Python, Spark Scala Developer

Infosys • Hyderabad

On-site
INR 1,000,000 - 1,500,000
Developer
Developer

GSB Solutions • Hyderabad

On-site
INR 1,000,000 - 1,500,000
Data Engineer
Data Engineer

Intact Green Services (india) • Bengaluru

On-site
INR 1,800,000 - 2,800,000
Industry-standard compensation
PySpark Developer / Senior Data Engineer
PySpark Developer / Senior Data Engineer

Alignity Solutions • Hyderabad

On-site
INR 1,200,000 - 1,800,000