Python & PySpark Developer
Job ID: 1484887
Location: Hyderabad
Experience: 5-7 Years
Openings: 1
Job Summary
We are looking for an experienced Python & PySpark Developer with strong expertise in PySpark, Spark SQL, ETL development, Delta Lake and Python-based data processing. The candidate should also have hands-on experience using GenAI/AI coding assistants to improve development, testing, debugging and documentation.
Primary Skills
- Python
- PySpark
- Spark SQL
- Hadoop
- Delta Lake
- ETL / Data Pipelines
- PySpark DataFrame
- SQL to PySpark Integration
- CSV / XML File Processing
- PyTest
- Python Debugging & Logging
Key Responsibilities
- Develop, maintain and enhance PySpark ETL jobs.
- Debug and optimize Spark DataFrame, Spark SQL and Delta Lake processing.
- Develop data pipelines for CSV, XML and other file-based ingestion.
- Implement reporting-date and business-date processing logic.
- Write and execute PyTest-based unit and integration tests.
- Troubleshoot production issues using Python logging, exception handling and debugging techniques.
- Work with Python packaging, runtime dependencies and code quality tools such as Pylint.
- Optimize PySpark transformations and Delta Lake read/write operations.
- Integrate SQL logic with PySpark-based data processing.
- Analyze existing code and support impact assessment, root-cause analysis and technical documentation.
GenAI / AI Skills
- Hands-on experience with GitHub Copilot, ChatGPT or similar AI coding assistants.
- Ability to use GenAI for Python, PySpark and SQL development.
- Experience using AI tools for ETL code analysis, debugging, impact assessment and documentation.
- Ability to use AI-assisted code review to identify bugs, performance issues, security risks and code-quality gaps.
- Experience generating or improving PyTest test cases using AI assistance.
- Understanding of prompt engineering for SQL analysis, Spark debugging, log analysis and documentation.
- Awareness of responsible AI usage, data privacy and validation of AI-generated outputs.
- Ability to use AI tools to understand legacy codebases and create technical documentation.
- Exposure to ML/AI, embeddings, vector search, RAG or LLM-based knowledge assistants is an advantage.
Required Experience
- 5-7 years of experience in Python/PySpark development.
- Strong hands-on experience with PySpark, Spark SQL and DataFrames.
- Experience with ETL pipelines and file ingestion.
- Good knowledge of Delta Lake and Spark processing.
- Experience with PyTest and production troubleshooting.
- Strong Python debugging and code-quality practices.
- Good communication and documentation skills.
Mandatory Skills
Python | PySpark | Spark SQL | Hadoop | Delta Lake | ETL | DataFrames | PyTest | SQL | GenAI/AI Coding Assistants