A leading technology firm in Hyderabad is looking for a Python Developer with strong proficiency in Python and PySpark to develop and maintain scalable data pipelines. The ideal candidate will have over 5 years of experience, particularly in solid data engineering practices. Key responsibilities include designing ETL processes and collaborating with teams to ensure data integrity. This is a great opportunity to work with cloud platforms and enhance your data handling skills.
Qualifications
5+ years of experience in Python and PySpark development.
Strong problem-solving and debugging skills.
Excellent communication and collaboration abilities.
Responsibilities
Develop and maintain scalable data pipelines using Python and PySpark.
Design and implement ETL processes.
Collaborate with cross-functional teams to understand data requirements.
Skills
Python programming
PySpark
Big Data technologies
SQL and databases
Data engineering best practices
REST APIs
Cloud computing environments
Machine learning libraries
Problem solving
Data quality
Job description
Qualifications and Experience
Strong proficiency in Python programming.
Hands‑on experience with PySpark and Apache Spark.
Knowledge of Big Data technologies (Hadoop, Hive, Kafka, etc.).
Experience with SQL and relational/non‑relational databases.
Familiarity with distributed computing and parallel processing.
Understanding of data engineering best practices.
Experience with REST APIs, JSON/XML, and data serialization.
Exposure to cloud computing environments.
5+ years of experience in Python and PySpark development.
Experience with data warehousing and data lakes.
Knowledge of machine learning libraries (e.g., MLlib) is a plus.
Strong problem‑solving and debugging skills.
Excellent communication and collaboration abilities.
Develop and maintain scalable data pipelines using Python and PySpark.
Design and implement ETL (Extract, Transform, Load) processes.
Optimize and troubleshoot existing PySpark applications for performance.
Collaborate with cross‑functional teams to understand data requirements.
Write clean, efficient, and well‑documented code.
Conduct code reviews and participate in design discussions.
Ensure data integrity and quality across the data lifecycle.
Integrate with cloud platforms like AWS, Azure, or GCP.
Implement data storage solutions and manage large‑scale datasets.