Position Title: Senior Data QA Engineer
Experience Level: 5+ Years
Job Overview
We are seeking an experienced Senior Data QA Engineer with 5+ years of hands-on experience in data testing, data warehousing, and modern data product validation. In this role, you will be responsible for ensuring the reliability, accuracy, completeness, and integrity of our enterprise data assets. You will work closely with Data Engineers, Data Architects, and Product Managers to build automated test frameworks, validate complex ETL/ELT pipelines, and assure high-quality data deliverables for business intelligence and data products.
Key Responsibilities
- Data Pipeline & Warehouse Testing:
- Design, develop, and execute comprehensive data test strategies, test plans, and automated test cases for large-scale ETL/ELT data pipelines.
- Validate data transformations, business logic, aggregation routines, and data models across Staging, Medallion architectures (Bronze/Silver/Gold), and Data Warehouses.
- Perform source-to-target data mapping validation, reconciliation, schema drift detection, and data completeness/integrity checks.
- Data Product Quality Assurance:
- Define and enforce data quality metrics (accuracy, completeness, freshness, consistency, uniqueness, and lineage) for customer-facing and internal data products.
- Validate end-to-end data flowsfrom ingestion sources (APIs, streaming, flat files) to downstream reporting, dashboards, and machine learning models.
- Conduct API and data interface testing to ensure seamless integration between data products and external/internal services.
- Collaboration & Governance:
- Partner with Data Architects and Engineers to identify root causes of data anomalies, performance bottlenecks, and pipeline failures.
- Establish data quality monitoring, alerting, and reporting dashboards to ensure proactive detection of data drift or corruption.
- Contribute to data governance, documentation, and compliance initiatives (e.g., data privacy, access controls, and auditing).
Qualifications & Skills
- Experience:
- 5+ years of dedicated experience in Data QA, Data Engineering, or Software Quality Assurance with a strong focus on data validation.
- Core Technical Expertise:
- Advanced SQL: Expertise in writing complex SQL queries, analytical functions, CTEs, joins, and optimization techniques for massive datasets.
- Data Warehousing: In-depth knowledge of modern cloud data platforms (e.g., Snowflake, Azure Synapse, AWS Redshift, BigQuery) and data modeling patterns (Star Schema, Snowflake Schema, Data Vault).
- Data Lakes & Formats: Experience testing structured and semi-structured file formats (Parquet, ORC, JSON, CSV) and cloud storage environments (AWS S3, Azure Data Lake Storage).
- Programming & Automation: Proficient in Python for writing automated test scripts and data comparison utilities.
- Data Product Knowledge:
- Solid understanding of the Data Product lifecycle, metadata management, data contracts, and source-to-target mapping documentation.
- Methodologies & Orchestration:
- Experience working with workflow orchestrators (e.g., Airflow, Dagster, Azure Data Factory).
Preferred Qualifications (Nice to Have)
- Experience with streaming data validation (e.g., Apache Kafka, Spark Streaming).
- Familiarity with data quality frameworks such as Great Expectations, dbt test, Soda, or custom Python/SQL test runners.
- Knowledge of Business Intelligence tools (e.g., Power BI, Tableau, Looker) to validate reporting datasets against underlying warehouse models.
- Exposure to cloud infrastructure and containerized testing environments (Docker).