Data Engineer (AWS, Spark)

Peregrine Advisors

Washington (District of Columbia)

Hybrid

USD 130,000 - 185,000

Full time

34 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Medical, dental, vision insurance (emp
401(k) with 100% match up to 4%
Unlimited PTO
Sponsored certifications

Job summary

Peregrine Advisors is seeking a data engineer to design and operate AWS-based pipelines that ingest, process, and store federal agency data. You will architect S3 data lakes with Iceberg, leverage Spark ETL on Glue/EMR, and build robust data quality, lineage, and governance.

You will work in a hybrid DC-area setting and partner with cross-functional teams to deliver auditable, scalable data solutions. You will need 4+ years of experience, Master’s in a relevant field, and familiarity with tools

Qualifications

  • 4+ years of data engineering experience.
  • Spark ETL on AWS (Glue, Amazon EMR) in Python and PySpark.
  • S3 data-lake design (Parquet, partitioning, lifecycle) feeding Apache Iceberg tables, Amazon Aurora PostgreSQL, and DynamoDB.
  • Event orchestration (Lambda, Step Functions, SQS/SNS) with secrets management and monitoring.
  • Data quality, validation, and lineage.
  • Infrastructure-as-code (CloudFormation or Terraform).
  • Federal information technology or high-volume data experience.
  • Familiarity with AI-assisted developer tooling.

Responsibilities

  • Build and maintain ingest-process-store pipelines on AWS.
  • Design and manage a data lake feeding platform stores via Iceberg, Aurora PostgreSQL, and DynamoDB.
  • Implement event orchestration and monitoring with CI/CD and secure access controls.
  • Ensure data quality, validation, and lineage across data flows.
  • Collaborate in a regulated environment ensuring accuracy and auditability.

Skills

Data engineering
AWS
ETL

Education

Master's degree in a relevant field

Tools

Glue
EMR
PySpark
Python
S3
Trino
Iceberg
Aurora PostgreSQL
DynamoDB
Lambda
Step Functions
CloudFormation
Terraform
SQS/SNS
Apache Ranger
DataStage

Job description

Peregrine Advisors is a firm founded on a simple conviction: the best solutions come from the people closest to the problem, given real ownership and the tools to deliver. We are a data and technology innovation hub and a Benefit Corporation working at the center of the federal government's mission to deliver for client stakeholders and the US public, looking for highly motivated contributors who thrive when trusted to own a hard problem and equipped to deliver the solution.

Your first assignment is to build the pipelines that move a federal agency's data from source to platform on Amazon Web Services (AWS): ingest, process, store, and keep it clean and trustworthy at scale.

The work is real, hard, and it matters. It is also where you start, not the shape of your career here: we hire people, not seats, and we move our best to where the hardest problems are.

  • Ingest-process-store pipelines on AWS: Spark-based extract, transform, and load (ETL) with Glue, Amazon EMR, Lambda, and Step Functions, in Python and PySpark.
  • A data lake that feeds the platform: S3 design (Parquet, partitioning, lifecycle) into Apache Iceberg tables, PostgreSQL on Amazon Aurora, and DynamoDB, with Trino for federated SQL across them, plus event orchestration, secrets and monitoring, and data quality, validation, and lineage built in.
  • In time, the firm itself: new capabilities, tools, and lines of business you help spin up.
  • Hands-on data engineering on AWS: Spark ETL (Glue, EMR), Python and PySpark, and S3 data-lake design feeding the platform stores.
  • The reliability craft around it: event orchestration, data quality and lineage, monitoring, and infrastructure-as-code.
  • The judgment to build in a regulated environment where accuracy and auditability are not optional.

Sole United States citizenship and the ability to obtain a Public Trust determination are required for this initial engagement. This is a hybrid role based in the Washington, DC metropolitan area, and it requires commuting into DC regularly. Everything else, the years, the certifications, the specific tools, we ask in the application and get into during the interview.

A high-performing team of developers, engineers, data scientists, architects, and strategists solving complex, real-world problems, with work that runs from strategy formulation to hands-on implementation. We develop people across roles and clients, with extensive onboarding and sponsored training and professional development. And Peregrine has been a Benefit Corporation from day one: public value is built into the work itself, not bolted on afterward. Work worth your best years.

As a Benefit Corporation, our commitment runs three ways: real, measurable value for our clients; government that works better for the public; and a team that makes everyone in it better.

We hire people who want to help build the firm, not just work at it.

Peregrine Advisors is an equal opportunity employer.

  • 4+ years of data engineering experience
  • Spark ETL on AWS (Glue, Amazon EMR) in Python and PySpark
  • S3 data-lake design (Parquet, partitioning, lifecycle) feeding Apache Iceberg tables, Amazon Aurora PostgreSQL, and DynamoDB
  • Event orchestration (Lambda, Step Functions, SQS/SNS) with secrets management and monitoring
  • Data quality, validation, and lineage
  • Infrastructure-as-code (CloudFormation or Terraform)
  • Basic proficiency in writing, PowerPoint, and Excel
  • Master's degree in a relevant field
  • Trino or comparable federated SQL across the lake and relational stores
  • Apache Ranger-governed access
  • Legacy ETL migration (for example DataStage)
  • Federal information technology or high-volume data experience
  • Familiarity with AI-assisted developer tooling
  • Medical, dental, and vision with the employee premium fully paid and half of dependent premiums; employer-paid life, accidental death, and short-term and long-term disability insurance; a 401(k) matched 100% up to 4% of salary, vesting immediately; unlimited paid time off; and sponsored professional certifications and continuing education.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Engineer (AWS, Spark)
Data Engineer (AWS, Spark)

Peregrine Advisors • Washington

Hybrid
USD 120,000 - 160,000
Health insurance
401(k) match
Unlimited PTO
+1
Data Engineer (AWS, Spark)
Data Engineer (AWS, Spark)

Uncover • Washington, Northern (KY)

Hybrid
USD 120,000 - 160,000
Hybrid work model
Sponsored training
Data Architect (AWS, Data Lake)
Data Architect (AWS, Data Lake)

Peregrine Advisors • Washington

Hybrid
USD 110,000 - 170,000
Data Analyst (SQL, AWS)
Data Analyst (SQL, AWS)

Peregrine Advisors • Washington

Hybrid
USD 75,000 - 110,000
401(k) matched
Paid time off
Professional certifications
+1
Senior Database Developer, AWS/PostgreSQL
Senior Database Developer, AWS/PostgreSQL

Peregrine Advisors • Washington

Hybrid
USD 140,000 - 180,000
Medical, dental, and vision covered
Employer-paid life and disability
401(k) with company match
+2
Data Scientist (Python, AWS)
Data Scientist (Python, AWS)

Peregrine Advisors • Washington

Hybrid
USD 120,000 - 180,000
Data Scientist (Python, AWS)
Data Scientist (Python, AWS)

Uncover • Washington, Northern (KY)

Hybrid
USD 90,000 - 130,000
Sponsored training
Python Developer (AWS, Cloud-Native)
Python Developer (AWS, Cloud-Native)

Peregrine Advisors • Washington

Hybrid
USD 120,000 - 150,000
Health benefits
401(k) matched
Unlimited PTO
+1
Data Architect (AWS, Data Lake)
Data Architect (AWS, Data Lake)

Uncover • Washington, Northern (KY)

Hybrid
USD 140,000 - 190,000
Hybrid work model in DC area
Benefit Corporation status
Data Analyst (SQL, AWS)
Data Analyst (SQL, AWS)

Uncover • Washington, Northern (KY)

Hybrid
USD 90,000 - 140,000
Hybrid work model
Professional development and extensive