Role Description
Required Skills & Experience
Around 5 years of professional experience in Data Engineering, Python development, or a related role.
Strong Hands-on Expertise In Python, Particularly For
- Data wrangling and transformation
- ETL/ELT development
- File and data processing
- Automation and scripting
- Data validation and cleansing
- Strong knowledge of at least one relational SQL database such as:
- PostgreSQL
- MySQL
- Oracle
- SQL Server
- Strong understanding of SQL, including joins, subqueries, CTEs, window functions, aggregations, and query optimization.
- Practical experience with AWS cloud services used in data engineering.
- Experience developing RESTful APIs using FastAPI.
- Good understanding of API concepts including HTTP methods, request/response handling, authentication, status codes, error handling, and API integration.
- Experience with Git/version control and standard software development practices.
- Good understanding of data engineering concepts, including data pipelines, ETL architecture, data quality, and data integration.
Good to Have
- Hands-on experience with PySpark and distributed data processing.
- Experience with NoSQL databases such as DynamoDB, or Cassandra.
- Experience with AWS services such as S3, Lambda, Glue, Athena, Redshift, EMR, RDS, Step Functions, EventBridge, and CloudWatch.
- Experience with workflow orchestration tools such as Apache Airflow or AWS Step Functions.
- Experience with Docker and containerized applications.
- Familiarity with CI/CD pipelines and DevOps practices.
- Knowledge of data warehousing and dimensional data modeling.
- Experience working with large-scale datasets and performance optimization.
- Familiarity with cloud-based data lake and data warehouse architectures.
Technical Skills
Skill Area Expected Knowledge
- Programming Python – Advanced
- Data Engineering ETL/ELT, data wrangling, transformation, validation
- Database SQL, Relational Databases
- Cloud AWS
- API Development FastAPI, REST APIs
- Big Data PySpark – Good to Have
- NoSQL MongoDB / DynamoDB / Cassandra – Good to Have
- Data Formats JSON, CSV, XML
- Version Control Git
- DevOps CI/CD, Docker – Good to Have