Qualifications
Primary skill: Python, Pyspark, Postgres, Data Pipeline QA
Notice Period: Immediate joiners
Experience: 5 to 7 Years
Responsibilities
- Develop and execute test strategies focused on data pipelines, including ingestion, transformation, storage, and retrieval.
- Perform data integrity validation by verifying schema consistency, data completeness, and correctness across distributed systems.
- Conduct Change Data Capture (CDC) testing to ensure seamless data updates and synchronization.
- Optimize and validate queries for performance and correctness using SQL, Trino, and Amazon Redshift.
- Implement automated data validation and regression tests to monitor for anomalies, data drift, and pipeline failures.
- Perform load testing and stress testing for large-scale data processing workflows.
- Identify bottlenecks and performance issues in ETL/ELT workflows and work with data engineers to optimize them.
- Ensure compliance with data governance standards, including security, retention, and access controls.
- Collaborate with engineers and reliability teams to define best practices for data quality and testing automation.
- Document test cases, defects, and resolutions, providing insights for continuous improvements.
Seniority level
Employment type
Job function
- Information Technology, Engineering, and Consulting
Industries
- IT Services and IT Consulting, Engineering Services, and Business Consulting and Services