We are looking for an experienced Data Solution Architect to design, develop, and optimize scalable data solutions within the AWS ecosystem. The ideal candidate will have strong expertise in AWS Glue, ETL pipeline development, Python, PySpark, and RDBMS/PostgreSQL, with the ability to architect robust data processing frameworks and ensure data quality, integrity, and performance. This role requires hands‑on experience in building cloud‑based data platforms, optimizing data workflows, and supporting enterprise‑scale data transformation initiatives.
Key Responsibilities
- Design and implement scalable ETL pipelines for structured and unstructured data processing.
- Develop and manage AWS Glue jobs, Data Catalogs, and scheduling workflows.
- Build efficient and optimized Python and PySpark scripts for large-scale data transformation.
- Work with PostgreSQL/Snowflake to write complex SQL queries, optimize performance, and manage database solutions.
- Ensure data quality, validation, integrity, and governance across data platforms.
- Manage and optimize AWS‑based data infrastructure, ensuring scalability, security, and reliability.
- Collaborate with cross‑functional teams to define data architecture and business requirements.
Required Skills & Evaluation Criteria
1. AWS Glue (Expert)
- Strong hands‑on experience in AWS Glue, ETL development, Data Catalog, and Job Scheduling.
- Ability to design, implement, and manage end‑to‑end ETL workflows using AWS Glue.
- Experience in optimizing Glue jobs for performance and scalability.
- Strong understanding of ETL architecture, data validation, and data integrity.
- Ability to design and implement robust, scalable, and fault‑tolerant ETL pipelines.
- Experience in handling large‑scale data ingestion and transformation.
- Proficiency in Python, OOP concepts, libraries, and exception handling.
- Ability to write clean, efficient, reusable, and scalable code.
- Experience with automation and scripting for data processing.
4. PySpark (Proficient)
- Strong knowledge of PySpark, DataFrames, RDDs, and Spark SQL.
- Ability to write and optimize PySpark scripts for distributed data processing.
- Experience in improving performance and handling large datasets.
- Hands‑on experience in PostgreSQL queries, DML operations, and database management.
- Strong expertise in Snowflake SQL, query optimization, and database design.
- Ability to write complex SQL queries and optimize database performance.
6. Data Platforms Management (Competent)
- Experience managing AWS ecosystem and cloud‑based data infrastructure.
- Understanding of security, governance, and platform optimization.
- Ability to ensure high availability and reliability of data platforms.