Data Engineer – Analytic Platform & Data Pipelines

3GIMBALS

United States

On-site

USD 135,000 - 185,000

Full time

9 days ago
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

3GIMBALS is seeking a Data Engineer to design, build, and maintain data pipelines and infrastructure powering our analytical platform. You’ll ingest, transform, and curate large volumes of data from diverse sources and build automated ETL/ELT workflows for downstream analytics, knowledge graph, and modeling teams.

The role emphasizes scaling with secure environments, governance, and data quality for high-velocity datasets.

Qualifications

  • 4+ years of data engineering experience building and operating production data pipelines.
  • Strong programming in Python and SQL (Scala/Java a plus).
  • Experience with distributed data processing frameworks (Spark, Dask, or similar).
  • Hands-on experience with workflow orchestration tools (Airflow, Dagster, Prefect).
  • Proficiency with relational and NoSQL databases (PostgreSQL, MongoDB, Elasticsearch).
  • Experience with cloud data platforms and services (AWS, Azure, or GCP).

Responsibilities

  • Design and build scalable batch and streaming pipelines to ingest structured and unstructured data from diverse sources.
  • Develop ETL/ELT workflows to normalize, enrich, and transform data into standardized schemas.
  • Build and maintain automated ingestion connectors for web, document, geospatial, and tabular data sources.
  • Manage data orchestration and scheduling using Airflow, Dagster, or Prefect.
  • Design and maintain data models, schemas, and storage across relational, NoSQL, and object stores.
  • Build and maintain data lakes/lakehouses and analysis-ready data marts.
  • Optimize partitioning, indexing, and query performance for large datasets.
  • Support entity resolution and data linking with knowledge graph and modeling teams.
  • Implement data validation, quality checks, and monitoring across pipelines.
  • Establish data lineage, cataloging, and metadata management.
  • Enforce data governance, provenance tracking, and source attribution for PAI/CAI data.
  • Document datasets, schemas, and pipeline logic for downstream consumers.

Skills

Python
SQL
Spark
Airflow
Docker
Kubernetes
Data Modeling
PostgreSQL
NoSQL

Education

Bachelor's degree in Computer Science
Data Engineering

Tools

Airflow
Dagster
Prefect
Docker
Kubernetes

Job description

Role Overview

3GIMBALS is seeking a Data Engineer to design, build, and maintain the data pipelines and infrastructure that power our unclassified PAI/CAI-based analytic platform. This role is responsible for ingesting, transforming, and curating large volumes of structured and unstructured data from diverse open and commercial sources; building resilient, automated ETL/ELT workflows; and ensuring data is high-quality, well-governed, and analysis-ready for the downstream analytics, knowledge graph, and modeling teams. The ideal candidate is comfortable working with messy, multi-source data at scale within secure development environments.

Role Overview

3GIMBALS is seeking a Data Engineer to design, build, and maintain the data pipelines and infrastructure that power our unclassified PAI/CAI-based analytic platform. This role is responsible for ingesting, transforming, and curating large volumes of structured and unstructured data from diverse open and commercial sources; building resilient, automated ETL/ELT workflows; and ensuring data is high-quality, well-governed, and analysis-ready for the downstream analytics, knowledge graph, and modeling teams. The ideal candidate is comfortable working with messy, multi-source data at scale within secure development environments.

Key Responsibilities
Data Pipeline Development & Ingestion
  • Design and build scalable batch and streaming pipelines to ingest structured and unstructured data from PAI/CAI sources, APIs, and third-party feeds
  • Develop ETL/ELT workflows to normalize, enrich, and transform heterogeneous data into standardized schemas
  • Build and maintain automated ingestion connectors for web, document, geospatial, and tabular data sources
  • Manage data orchestration and scheduling using tools such as Airflow, Dagster, or Prefect
Data Modeling & Storage
  • Design and maintain data models, schemas, and storage layers across relational, NoSQL, and object stores
  • Build and maintain data lakes/lakehouses and curated, analysis-ready data marts
  • Optimize partitioning, indexing, and query performance for large datasets
  • Support entity resolution and data linking in coordination with the knowledge graph and modeling teams
Data Quality, Governance & Lineage
  • Implement data validation, quality checks, and monitoring across pipelines
  • Establish data lineage, cataloging, and metadata management
  • Enforce data governance, provenance tracking, and source attribution appropriate for PAI/CAI data
  • Document datasets, schemas, and pipeline logic for downstream consumers
Security & Compliance
  • Ensure pipelines and data stores meet security requirements for operation in sensitive environments
  • Implement encryption, access control, and secure data-handling practices
  • Support Authority to Operate (ATO) processes and compliance frameworks
Required Qualifications
Technical Expertise
  • 4+ years of data engineering experience building and operating production data pipelines
  • Strong programming skills in Python and SQL (Scala or Java a plus)
  • Experience with distributed data processing frameworks (Spark, Dask, or similar)
  • Hands-on experience with workflow orchestration tools (Airflow, Dagster, Prefect)
  • Proficiency with relational and NoSQL databases (PostgreSQL, MongoDB, Elasticsearch, etc.)
  • Experience with cloud data platforms and services (AWS, Azure, or GCP)
Data & Infrastructure
  • Experience designing data models, warehouses, and lakehouse architectures
  • Familiarity with data formats and serialization (Parquet, Avro, JSON, GeoJSON)
  • Understanding of data quality, lineage, and governance practices
  • Experience with containerization (Docker) and CI/CD for data workflows
Domain Knowledge
  • Experience working with large-scale, heterogeneous, or open-source datasets
  • Understanding of data provenance and source-attribution requirements
Preferred Qualifications
  • Active security clearance or ability to obtain one
  • Experience in government, defense, or intelligence contracting environments
  • Familiarity with PAI/CAI (publicly and commercially available information) data sources
  • Experience with geospatial data processing (PostGIS, GDAL, or similar)
  • Knowledge of graph data structures and preparing data for knowledge graphs
  • Experience with streaming platforms (Kafka, Kinesis)
  • Familiarity with federal compliance frameworks (FedRAMP, FISMA, NIST 800-53)
Technical Environment
  • Languages: Python, SQL (Scala/Java a plus)
  • Processing: Spark, Airflow/Dagster/Prefect, streaming frameworks
  • Storage: PostgreSQL, Elasticsearch, object storage / data lake, Parquet
  • Infrastructure: Docker, Kubernetes, cloud platforms (AWS GovCloud, Azure Government)
  • Security: Encryption at rest and in transit, RBAC, secure data handling

This role is central to the platform: the data engineering team delivers the clean, trustworthy, well-documented data that every analytic, knowledge graph, and risk-modeling capability depends on.

Salary: $135000 - $185000 per year

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Data Engineer – Analytic Platform & Data Pipelines
Data Engineer – Analytic Platform & Data Pipelines

3GIMBALS • Virginia (MN)

On-site
USD 110,000 - 160,000
Data Engineer: Scalable, Governed Pipelines for Analytics
Data Engineer: Scalable, Governed Pipelines for Analytics

3GIMBALS • United States

On-site
USD 135,000 - 185,000
Data Engineer: Scalable Pipelines for Analytics Platform
Data Engineer: Scalable Pipelines for Analytics Platform

3GIMBALS • Virginia (MN)

On-site
USD 110,000 - 160,000
Data Engineer
Data Engineer

Mondo • Baltimore (MD)

Hybrid
USD 100,000 - 150,000
Medical
Dental
Vision
+4
Data Engineer
Data Engineer

Strategic Innovation Group • Arlington (VA)

On-site
USD 120,000 - 170,000
Health insurance
401(k) with match
PTO
+4
Data Engineer
Data Engineer

Strategic Innovation Group, LLC • Arlington (VA)

On-site
USD 150,000 - 185,000
Health insurance
Dental insurance
Vision insurance
+4
Data Engineer - III
Data Engineer - III

Compunnel, Inc. • San Francisco (CA)

On-site
USD 120,000 - 150,000
Senior Data Engineer
Senior Data Engineer

ANAUTICS INC • Oklahoma City (OK)

On-site
USD 90,000 - 120,000
Data Engineer
Data Engineer

Prodigy Resources • Denver (CO)

On-site
USD 110,000 - 170,000
Senior Data Engineer
Senior Data Engineer

Madison-Davis, LLC • Chicago (IL)

On-site
USD 130,000 - 180,000