An application made for this job — a tailored resume and cover letter that speak straight to the posting.
Jobtailor is seeking a Senior Data Engineer in Maryland to design, develop, and maintain robust data pipelines from external feeds into our data lake and downstream zones.
You will implement production-grade ETL workflows with AWS Glue, PySpark, Lambda, and MWAA, loading data into S3, Redshift, Oracle, and other stores while ensuring data quality and traceability.
• Design, develop, ingest, and maintain well-architected data pipelines that retrieve data from external feeds (APIs, SFTP, HTTPS, FTP, web scraping, Direct Connect) and internal agency sources into the data lake landing zone and downstream curated zones.
• Develop production-grade ETL workflows using AWS Glue, PySpark, Python, Lambda, and EMR, integrated with a shared ETL common library and orchestrated via Amazon Managed Workflows for Apache Airflow (MWAA).
• Load data accurately and optimally into S3 zones (Parquet, ORC, Iceberg), relational datastores (PostgreSQL, Redshift, Oracle), NoSQL databases, and knowledge bases/vector stores, preventing duplicate loads and maintaining data integrity and traceability across all lifecycle stages.
• Implement schema enforcement, XSD validation, data quality checks, error handling, and automated SNS notifications; ensure all production jobs populate ETL Load Reports and Gap Reports through static and dynamic ETL metadata.
• Develop semantic-layer objects (tables, views, materialized views) that ensure complete data coverage, optimized query performance, and consistent application of business logic.
• Develop XML parsing/shredding logic for high-volume regulatory filings using Glue PySpark, supporting schema evolution and batch processing per program standards.
• Design pipelines with query performance in mind and support rollback, reload, and date-range reprocessing capabilities without manual intervention.
• Support self-service ETL development by other agency teams through standardized, reusable components aligned with program standards, and facilitate the transition of externally developed ETL jobs into the Data Engineering team's production support.
• Create and maintain required engineering artifacts, including business requirements, ETL design documents, mapping documents, data models, data dictionaries, deployment references, operations and maintenance guides, and test plans.
• Deploy code through automated CI/CD pipelines using CloudFormation templates, following agency release, security, and governance processes.
• Provide operational support for production jobs, including rapid identification and resolution of failed jobs and performance issues; participate in on-call/after-hours support for production outages and emergencies as part of a team rotation.
• Collaborate with Data Officers, Data Stewards, SMEs, data providers, and IV&V teams to understand requirements and deliver user-accepted solutions; engage closely with the Product Owner and cross-functional teams to provide timely updates and resolve issues.
• Leverage AI-assisted development tools to accelerate coding, optimize workflows, and enhance code quality while adhering to security and performance standards.
• Work in Agile teams; drive iterative delivery, joint problem-solving, and continuous improvement, including participation in sprint planning and program increment planning.
Requirements
Core Competencies
Demonstrates expertise in designing and developing data pipelines using AWS Glue, PySpark, and Python, with a strong focus on ETL workflows and data integrity. Proficient in collaborating with cross-functional teams to deliver high-quality data solutions while adhering to Agile methodologies.
Highest-signal resume keywords
ATS Optimization Keywords
Hard Skills
Soft Skills
Industry Keywords
Tools & Technologies