Data Engineer Role

Peregrine Advisors LLC

Washington (District of Columbia)

On-site

USD 115,000 - 165,000

Full time

5 hours ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Peregrine Advisors LLC is seeking a Data Engineer to design and operate end-to-end data pipelines that move data from sources to people, applications, analyses, and models. You will build ingestion, transformation, and storage components, ensuring accuracy, security, and traceability at scale.

You will collaborate with data scientists, engineers, and governance professionals to maintain data contracts and trustworthy data products across cloud and on‑prem environments.

Qualifications

  • Working foundation in programming and query languages used for data engineering (Python/SQL).
  • Experience with data ingestion, transformation, storage, and data-model design.
  • Understanding batch, streaming, and event-driven processing and data governance concepts.
  • Ability to work with orchestration, testing, deployment, monitoring, and documentation.

Responsibilities

  • Build and operate ingestion and transformation pipelines for diverse data sources.
  • Develop and maintain cloud/on-prem data platforms and serving layers.
  • Ensure data quality, lineage, metadata, and access controls to keep data trustworthy.
  • Collaborate with engineers, analysts, data scientists, and governance specialists.
  • Document data contracts, tradeoffs, and data-product operational aspects.

Skills

Python
SQL
Data engineering
Distributed data systems
ETL/ELT concepts

Tools

Airflow
dbt
Kafka
Spark

Job description

Description
The work

Data Engineers build and operate the systems that move data from its sources to the people, applications, analyses, and models that depend on it. They ingest data, transform it, organize it for use, and keep it accurate, secure, traceable, and available at scale. Their work turns fragmented files, documents, databases, application programming interfaces (APIs), and event streams into reliable data products.

Artificial intelligence (AI) is one important consumer of that work, alongside reporting, visualization, analytics, software, and operational systems. Some openings may involve preparing dependable data for machine learning, document retrieval, or AI evaluation. The center of the Role remains dependable data engineering from source to use.

What you may build
  • Ingestion and transformation pipelines for batch, streaming, and event-driven data from APIs, databases, files, documents, object stores, messaging systems, and operational platforms.
  • Cloud and on-premises data platforms, including databases, data lakes, warehouses, lakehouses, and serving layers for reporting, visualization, software, analytics, and other operational uses.
  • Quality, validation, metadata, lineage, provenance, and access-control capabilities that make data trustworthy, explain how it changed, and keep its use within approved boundaries.
  • The operational layer around data products: orchestration, testing, monitoring, backfills, replay, recovery, retention and deletion implementation, performance and cost tuning, infrastructure as code, and technical documentation.
  • Where the work requires it, versioned feature, training, testing, or evaluation datasets; document and retrieval-index pipelines; or governed telemetry and feedback data that support machine learning and generative AI systems.
Who you are

You care whether data arrives, but also whether it is complete, timely, understood, authorized, and fit for use. You trace failures across sources, transformations, storage, and serving layers, and you improve recurring processes instead of working around them.

You collaborate well with source-system owners, software engineers, analysts, data scientists, AI Engineers, Machine Learning Engineers, security and governance specialists, and Development, Security, and Operations (DevSecOps) Engineers. You make data contracts and tradeoffs clear, distinguish a data problem from a model or application problem, and prefer ownership of an outcome to a narrowly assigned task.

What you bring
  • A working foundation in programming and query languages used for data engineering. Python and Structured Query Language (SQL) are common, but the specific stack varies by opening.
  • Experience or strong grounding in data ingestion, transformation, storage, schema and data-model design, and the performance characteristics of distributed data systems.
  • An understanding of batch, streaming, and event-driven processing, together with orchestration, testing, deployment, monitoring, recovery, and documentation.
  • Practical experience with data quality, metadata, lineage, provenance, versioning, access controls, and secure data handling.
  • The judgment to work in environments where accuracy, privacy, security, traceability, reproducibility, resilience, performance, and cost matter.
What future openings may require

Each opening will identify the experience, platform, tooling, data, performance, security, location, work-authorization, citizenship, suitability, clearance, and domain knowledge the work requires. Those requirements will vary, and no candidate is expected to cover every specialization.

An opening may emphasize extract, transform, and load (ETL), extract, load, and transform (ELT), batch or stream processing, data modeling, a warehouse or lakehouse, document extraction, a feature store, machine-learning data, search or vector indexing, retrieval-augmented generation data preparation, telemetry and feedback pipelines, data-governance implementation, or platform operations.

Specific openings may name Amazon Web Services, Microsoft Azure, Google Cloud, Spark, Airflow, Kafka, Flink, dbt, relational or nonrelational databases, data warehouses, lakehouses, search platforms, vector databases, Linux, container platforms, or infrastructure-as-code tools. OPEN Data Jobs will state those requirements with the opening rather than treat every technology in this Role description as universal.

About OPEN Data Jobs

OPEN Data Jobs connects artificial intelligence, data, and software professionals with federal-sector opportunities. OPEN Data Jobs is a division of Peregrine Advisors Benefit, Inc.

Benefits

Compensation, benefits, work location, and employment terms are set for each specific opening and will be stated with that opening

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Data Scientist Role
Data Scientist Role

Peregrine Advisors LLC • Washington

On-site
USD 120,000 - 180,000
Data Program Manager Role
Data Program Manager Role

Peregrine Advisors LLC • Washington

On-site
USD 120,000 - 150,000
AI Engineer Role
AI Engineer Role

Peregrine Advisors LLC • Washington

On-site
USD 120,000 - 190,000
Data Engineer
Data Engineer

Prodigy Resources • Denver (CO)

On-site
USD 110,000 - 170,000
Data Analyst Role
Data Analyst Role

Peregrine Advisors LLC • Washington

On-site
USD 65,000 - 95,000
Generative AI Engineer Role
Generative AI Engineer Role

Peregrine Advisors LLC • Washington

On-site
USD 140,000 - 210,000
Data Platform Engineer
Data Platform Engineer

JACK Main • Cleveland (OH)

On-site
USD 80,000 - 110,000
Sr. Data Engineer
Sr. Data Engineer

Save A Lot • St. Ann (MO)

On-site
USD 110,000 - 160,000
401K match
Paid Time Off
Medical insurance
+6
Data Strategist Role
Data Strategist Role

Peregrine Advisors LLC • Washington

On-site
USD 120,000 - 180,000
Data Product Owner Role
Data Product Owner Role

Peregrine Advisors LLC • Washington

On-site
USD 120,000 - 160,000