MANTECH seeks a motivated, career and customer-oriented Lead Data Engineer to join our Enterprise Data, AI, and Automation team in Herndon, VA. This is a hybrid position, requiring 2-3 days a week onsite.
Responsibilities Include But Are Not Limited To
- Enterprise Pipeline Engineering: Design, build, and maintain secure, scalable ETL/ELT pipelines integrating disparate enterprise business systems and API connections into a unified enterprise data model.
- Data Warehousing & Modeling: Develop/update data warehouse schema to align with our evolving business requirements for reporting, analytics, and automation.
- Tooling Strategy: Act as a subject matter expert in the selection and implementation of next-generation analytics platforms and data engineering tools.
- Lakehouse & Semantic Architecture: Drive the migration toward a modern data lakehouse and assist analytics engineers with implementation of a universal semantic data model.
- Data Quality & Observability Infrastructure: Implement automated data validation, lineage tracking, and end-to-end observability to ensure high data fidelity and pipeline reliability across the enterprise platform.
- Proactive Alerting & Executive Intelligence: Support analytics engineers to implement automated orchestration logic and alerting triggers that notify business leaders in real time when key performance metrics cross predefined limits.
- Team Mentorship & Support: Guide teammates on modern data engineering practices and architect data flows optimized for consumption by machine learning models and AI applications.
Minimum Qualifications
- Bachelor’s degree in Computer Science, Information Systems, Data Engineering, or a related field with 7+ years of data engineering experience, including at least 2+ years leading engineering initiatives or architectural design for enterprise systems.
- Proven experience integrating common enterprise systems via database connectors or APIs.
- Hands-on experience working with a data lakehouse architecture and proven knowledge of tradeoffs associated with data modeling approaches.
- Mastery in writing complex, optimized SQL queries and managing relational database schemas.
- Experience with modern cloud data platforms, transformation tools, and code-driven orchestration tools (e.g., Apache Airflow, Dagster, Prefect).
- Experience with common software tools (e.g., Azure DevOps, GitHub) for CI/CD and version control of data infrastructure.
Preferred Qualifications
- Proficiency in Python and PySpark for data engineering, data manipulation, API consumption, and automation scripts.
- Experience using Informatica for API and database extraction.
- Experience implementing automated data quality, monitoring, and lineage tooling (e.g., Monte Carlo, Great Expectations, Soda, or Databricks Unity Catalog).
- Experience building data pipelines for machine learning, natural language processing, or vector database/RAG workflows.
- Experience with Docker and Infrastructure-as-Code (Terraform or CloudFormation).
Clearance Requirements
Physical Requirements
- Must be able to remain in a stationary position 50%.
- Needs to occasionally move about inside the office to access file cabinets, office machinery, etc.
- Frequently communicates with co-workers, management, and customers, which may involve delivering presentations. Must be able to exchange accurate information in these situations.