ScolerTec is seeking multiple Lead Data Engineers to support a large-scale cloud data modernization and governance program. This role will lead the design, development, testing, and maintenance of secure, scalable data engineering components and platform services supporting operational data, reporting, analytics, and AI/ML use cases. The Lead Data Engineer will work closely with architecture, governance, migration, and DevSecOps teams to deliver high-quality, auditable, and performant cloud-based data solutions using AWS and Databricks.
Key Responsibilities
- Lead development of scalable cloud-based data platforms supporting data lake, lakehouse, and data warehouse architectures.
- Design and develop enterprise ETL/ELT and ingestion pipelines for batch and near-real-time workloads.
- Build and optimize data engineering solutions using Databricks, Spark/PySpark, Delta Lake, Python, and SQL.
- Implement secure and auditable integrations across legacy, cloud, and external systems.
- Lead implementation of operational databases, document storage, schemas, APIs, and data services.
- Provide technical direction and oversight to Data Engineers, including design and code reviews.
- Implement Infrastructure as Code for database and platform provisioning and configuration.
- Implement data quality checks, validation, error handling, metadata, lineage, and governance-aligned structures.
- Optimize pipeline performance, Spark workloads, query performance, scalability, and cloud cost.
- Integrate pipelines with Government-provided CI/CD processes.
- Support incident resolution, RCA, migration, archival, failover testing, and DR activities.
- Maintain technical documentation, architecture artifacts, runbooks, and audit-ready deliverables.
Qualifications
- 7+ years of specialized data engineering experience, including leadership of enterprise-scale engineering efforts.
- Strong hands-on experience with AWS data and compute services.
- Strong hands-on production experience with Databricks, Spark/PySpark, and Delta Lake/lakehouse architectures.
- Strong proficiency in SQL, Python, ETL/ELT, and data transformation.
- Hands-on experience building enterprise-scale batch and near-real-time data pipelines.
- Experience with AWS data services such as S3, Glue, Redshift, Kinesis, Lambda, Athena, or related services.
- Experience with orchestration tools and streaming/near-real-time ingestion.
- Experience with CI/CD, secure software engineering, and Infrastructure as Code.
- Experience with data governance, metadata, lineage, data quality, and classification standards.
- Experience leading and mentoring Data Engineers.
- Federal or regulated-environment experience preferred.
Preferred
- Databricks experience with Unity Catalog, Workflows, Auto Loader, SQL Warehouses, Delta Live Tables/Lakeflow, or governed Gold-layer datasets.
- Informatica or another enterprise integration/ETL platform.
- Terraform or equivalent Infrastructure as Code.
- Kafka or Kinesis streaming experience.
Preferred Certifications
- AWS Certified Solutions Architect – Professional
- AWS Certified Data Analytics – Specialty