Central Strategies is seeking an AI/ML Data Engineer to support the U.S. Coast Guard. This position supports the Coast Guard’s Digital Transformation Strategy by designing, building, deploying, and scaling secure data platforms and production-grade data pipelines that enable Artificial Intelligence (AI), Machine Learning (ML), advanced analytics, and data-driven automation.
The AI/ML Data Engineer will work closely with data scientists, ML engineers, cloud engineers, enterprise architects, cybersecurity professionals, automation engineers, mission stakeholders, and program leadership to transform complex data from multiple sources into trusted, reusable, and high-performance data products for AI/ML and analytics solutions.
This role supports an enterprise automation and data modernization program focused on lakehouse architecture, cloud-native data integration, MLOps, data governance, workflow modernization, and reliable delivery of mission-ready data capabilities across the U.S. Coast Guard.
Work Location: Hybrid in the Washington, DC metropolitan area.
Key Responsibilities
- Design, build, test, deploy, and maintain scalable data platforms that support Coast Guard AI/ML, analytics, and mission operations.
- Engineer reusable data services and curated data products for predictive analytics, generative AI, natural language processing, anomaly detection, and decision-support solutions.
- Translate AI/ML use cases into reliable data architecture, storage, processing, and integration requirements.
- Support the transition of AI-enabled minimum viable products (MVPs) and pilots into secure, production-scale capabilities.
- Develop and optimize batch and streaming ETL/ELT pipelines for structured, semi-structured, and unstructured data.
- Build automated ingestion, transformation, validation, and publishing workflows across multiple Coast Guard systems and repositories.
- Implement orchestration, monitoring, logging, error handling, and recovery mechanisms for reliable production data pipelines.
- Apply software engineering best practices, including modular design, version control, automated testing, documentation, and code review.
Databricks & Lakehouse Engineering
- Design and implement Databricks lakehouse solutions using Apache Spark, Delta Lake, Unity Catalog, notebooks, workflows, and SQL warehouses.
- Develop scalable data models and medallion architectures that support analytics, feature engineering, model training, and model inference.
- Optimize Spark workloads, data partitioning, storage layouts, and compute utilization for performance, reliability, and cost efficiency.
- Integrate Databricks with cloud storage, databases, APIs, event streams, business applications, and enterprise data services.
- Implement governed data access, lineage, metadata management, and reusable data-sharing patterns.
ML Ops & Production Integration
- Build and maintain data and feature pipelines that support model development, training, evaluation, deployment, monitoring, and retraining.
- Collaborate with data scientists and ML engineers to operationalize models, RAG pipelines, vector search, model endpoints, and generative AI applications.
- Integrate CI/CD, infrastructure as code, automated testing, containerization, and release controls into AI/ML data engineering workflows.
Data Quality, Security & Governance
- Implement automated data-quality rules, reconciliation, observability, and monitoring to ensure accurate, complete, and trusted data.
- Apply Federal security, privacy, access-control, audit, retention, and governance requirements throughout the data lifecycle.
- Diagnose data, pipeline, and platform performance issues and implement durable corrective actions.
- Participate in Agile planning, backlog refinement, sprint reviews, technical demonstrations, release activities, and operational support.
- Collaborate with product owners, architects, developers, cybersecurity teams, data scientists, and mission stakeholders to translate requirements into production data capabilities.
- Create technical documentation, data dictionaries, architecture artifacts, runbooks, and knowledge-transfer materials.
Required Qualifications
- Bachelor’s degree in Computer Science, Data Engineering, Software Engineering, Information Systems, Artificial Intelligence, Engineering, or related field.
- 7+ years of experience in data engineering, software engineering, cloud data platforms, data architecture, or related technical disciplines.
- Experience designing, developing, deploying, and operating production-grade batch and streaming data pipelines.
- Strong proficiency with:
- Python and PySpark
- SQL
- Apache Spark and distributed data processing frameworks
- Experience designing and implementing:
- Batch and real-time streaming solutions
- Dimensional, relational, and lakehouse data models
- Workflow orchestration, monitoring, and automated recovery
- Experience working with large datasets, complex source systems, and high-volume data environments.
- Hands‑on experience with Databricks, including notebooks, workflows, Delta Lake, Unity Catalog, and performance optimization.
- Experience building data and feature pipelines that support machine learning, generative AI, RAG, vector search, and model-serving workloads.
- Experience with data integration technologies such as Kafka, Airflow, AWS Glue, APIs, message queues, or change data capture.
- Experience engineering data solutions in AWS, Azure, or GovCloud environments.
- Experience integrating modern data engineering practices with:
- CI/CD pipelines and automated testing
- Infrastructure as code, such as Terraform or CloudFormation
- Git-based version control and code review
- Containers and orchestration platforms, such as Docker or Kubernetes
- MLOps tools and model lifecycle platforms
- Strong communication skills with experience documenting technical solutions and briefing technical and non-technical stakeholders.
- Ability to obtain and maintain a DHS Public Trust.
Preferred Qualifications
- Experience supporting DHS, USCG, DoD, or other Federal agencies; experience implementing governed data platforms in regulated environments is preferred.
Desired Certifications
- Databricks Certified Data Engineer Associate is a plus.