A leading technology solutions provider in Reston, VA is seeking a skilled data engineer. The role involves building and maintaining ETL pipelines using Python and PySpark on AWS Glue, as well as orchestrating workflows with AWS services. Candidates must have proven experience with AWS tools like Lambda, SNS, SQS, and Redshift, alongside solid SQL skills. This position promotes teamwork and data quality assurance while contributing to efficient data management solutions.
Qualifications
Strong experience with Python and PySpark for large-scale data processing.
Solid SQL skills and familiarity with data modeling and query optimization.
Experience with ETL best practices, data quality checks, and monitoring/alerting.
Familiarity with version control (Git) and basic DevOps/CI-CD workflows.
Responsibilities
Build and maintain ETL pipelines using Python and PySpark.
Orchestrate workflows with AWS Step Functions and serverless components.
Implement messaging and event-driven patterns using AWS SNS and SQS.
Design and optimize data storage and querying in Amazon Redshift.
Write performant SQL for data transformations, validation, and reporting.
Ensure data quality and monitoring for pipelines.
Collaborate with data consumers and stakeholders.
Skills
Python
PySpark
AWS services
SQL
ETL best practices
Git
Job description
A leading technology solutions provider in Reston, VA is seeking a skilled data engineer. The role involves building and maintaining ETL pipelines using Python and PySpark on AWS Glue, as well as orchestrating workflows with AWS services. Candidates must have proven experience with AWS tools like Lambda, SNS, SQS, and Redshift, alongside solid SQL skills. This position promotes teamwork and data quality assurance while contributing to efficient data management solutions.