Senior Data Engineer

Voreas Laboratories Inc.

Washington, Northern (District of Columbia, KY)

Hybrid

USD 140,000 - 190,000

Full time

22 hours ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Fully remote within the US
End-to-end ownership
Flexible hours

Job summary

Voreas Laboratories Inc. is seeking a Senior Data Engineer to own production-grade data pipelines in a US-remote setup. You will work on lake data, detections, and evidence packaging, ensuring correctness at scale.

Candidates should have hands-on experience with Spark, Iceberg/Delta Lake, containerization, workflow orchestration (Argo Workflows), and event-time correctness. US citizenship required.

Qualifications

  • Spark in production at high volume with tuning experience.
  • Open table format in production: Iceberg or Delta Lake.
  • Containerized someone else's code and scheduled production runs.
  • Workflow orchestration in production (Argo Workflows, Airflow, Dagster, Step Functions).
  • Event-time correctness; understand processing vs event time.
  • Ability to take something that worked in one env and run it autonomously.
  • Comfort with correctness as a matter of degree.
  • Must be a US citizen.

Responsibilities

  • Write Spark jobs to slice lake telemetry down to detector-ready granularity.
  • Containerize detections and pin dependencies for reproducible runs.
  • Orchestrate end-to-end workflows to run at scale and recover from partial failures.
  • Write findings back to canonical tables with provenance (detector, version, window, data).
  • Ensure data quality by validating input schemas and outputs.

Skills

Spark production tuning
Event-time correctness
Production deployment
Remote work experience
US citizenship

Tools

Apache Iceberg
Delta Lake
Amazon EKS
Argo Workflows
Argo CD
Crossplane
Containerization

Job description

Engineering – Senior Data Engineer – Atlanta or DC Area Preferred – Full-Time

DNS and netflow telemetry lands in a data lake in our own AWS account, detection pipelines score it, and the findings are written into an intelligence graph. Security researchers design the detections and often build them. This team builds everything around them: the pipelines that reduce lake data to what a detector needs, the packaging that lets it run on our platform, the orchestration that runs it on schedule at full volume, and the path that writes what it found back to the lake as evidence.

Our stack is Amazon EKS, Spark, Apache Iceberg and Delta Lake, Argo Workflows, Argo CD, Crossplane, and Python and Scala. We do not expect you to have used all of it. We do expect that you have worked in each of these areas, in whatever tools you used at the time: continuous integration and delivery, infrastructure as code, distributed data processing, container orchestration, a data lake on an open table format, and Python.

This role gets detections into production and keeps them there. A researcher hands you a detection, sometimes as working code, sometimes as a written technique. Everything between the data lake and a finding recorded as evidence is yours. Most of the difficulty is correctness under conditions a researcher's environment never faces: the volume is far larger, the data has gaps and skew, the job will be re-run, and someone will eventually need to explain why a finding says what it says.

Location
Atlanta or DC Area Preferred

Committment
Full-Time

Work style
Remote

Responsibilities
  • Reduction. Detectors do not run against the whole lake. Write the Spark jobs that cut telemetry at our volume down to the slice a detector needs, at the grain it expects.
  • Packaging. Containerize the detection so it runs the same way every time, with its dependencies pinned and its interface fixed.
  • Orchestration. Build the workflow that runs it on schedule at full volume, survives partial failure, and can be re-run over a past window without producing duplicate findings.
  • Write-back. Write what it found into the canonical tables as evidence, with the provenance that makes a finding defensible: which detector, which version, which window, and what data it was based on.
  • Data quality in both directions. Verify that incoming data conforms to the schema we specified, and catch output that does not look like what the detection predicted.
Required Qualifications
  • Spark in production at volume, with real tuning experience. Skew, partitioning, shuffle behavior, join strategy.
  • An open table format in production: merges, incremental processing, compaction and table maintenance. We run both Apache Iceberg and Delta Lake. Either one is fine.
  • You have containerized someone else's code and run it on a schedule in production. Dependency pinning, resource limits, and what happens when it fails halfway through.
  • Workflow orchestration in production, in any tool. Argo Workflows, Airflow, Dagster, Step Functions. We use Argo Workflows and will not test you on it.
  • Event-time correctness. You know why processing time and event time differ and what breaks when they are conflated.
  • You have taken something that worked in one person's environment and made it run without supervision. This is most of the job and we will ask about a specific case.
  • Comfort working where correctness is a matter of degree rather than a passing test.
  • Must be a US citizen
Preferred Qualifications
  • Our specific stack: Amazon EKS, Argo Workflows, Argo CD, Crossplane.
  • Scala. Not required, though some of what you maintain is written in it.
  • High-volume network telemetry.
Why Join Us
  • Fully remote within the United States, with flexible working hours
  • End-to-end ownership: everything between the data lake and a recorded finding
  • A modern stack and an engineering culture that values automation and best practices
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Technical Lead, Data Engineering
Technical Lead, Data Engineering

Voreas Laboratories Inc. • Washington, Northern (KY)

Hybrid
USD 180,000 - 230,000
Fully remote
Flexible hours
Design ownership
+1
Senior Software Platform Engineer
Senior Software Platform Engineer

Voreas Laboratories Inc. • Northern (KY)

Hybrid
USD 130,000 - 190,000
Data Engineer
Data Engineer

detections-ai • California (MO)

Remote
USD 130,000 - 180,000
Remote-first equity
Equity
Remote-first culture
Data Engineering Manager
Data Engineering Manager

Hidden Jobs • United States

Remote
USD 187,000 - 253,000
Data Engineer
Data Engineer

7Seventy • Northern (KY)

On-site
USD 90,000 - 130,000
Lead Data Engineer
Lead Data Engineer

Norfolk Southern Corp • Atlanta (GA)

Hybrid
USD 150,000 - 230,000
Security Engineer, Detection & Response
Security Engineer, Detection & Response

United States Digital Space LLC • New York (NY), Washington

On-site
USD 238,000 - 297,000
Health, dental & vision coverage
Retirement benefits
Learning & development stipend
+2
Data Engineer
Data Engineer

SpaceCoast AV Consultants • Town of Florida (NY)

On-site
USD 110,000 - 140,000
Health, dental, and vision insurance
401(k) with company matching
Generous paid time off and parental leave
+1
Detection Engineer, Security Operations & Telemetry
Detection Engineer, Security Operations & Telemetry

Saronic • Austin (TX)

On-site
Confidential
Data Engineer (in person)
Data Engineer (in person)

SEP • Westfield (IN)

On-site
USD 90,000 - 110,000
Flexible work schedules
Opportunities to learn and develop
Community of friendly peers
+1