Data Platform Engineer

Nebula

El Segundo (CA)

Hybrid

USD 100,000 - 250,000

Full time

18 hours ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Equity participation
Health, dental and vision insurance
401(k) with company matching
Bonus for living within 5 miles of ElD

Job summary

Nebula is seeking a data engineer to build and operate the production data platform that stores, models and streams knowledge about customers, machines and the company itself. You will own pipelines, entity model, provenance and time-series infrastructure, collaborating with AI systems daily.

You will work on-site in El Segundo, with equity and strong compensation. This role demands hands-on experience with data pipelines, SQL, Python, and modern data stores, and a passion for scalable, reliable

Qualifications

  • Degree in computer science, data engineering, information systems or related field or equivalent experience
  • Hands-on production data pipelines with idempotency, replay, backfills and schema evolution
  • Information modeling and entity resolution with canonical identifiers and reversible merges
  • Fluent with document and relational stores; SQL and Python; able to explain slow queries
  • Judgment to trace wrong numbers through pipelines rather than altering outputs
  • Ready to build first version of a system and hand off to next engineer
  • Bonus: knowledge graphs, taxonomy design, geospatial work or high-rate time-series

Responsibilities

  • Build and manage the data estate: schemas, indexes, migrations, query performance, retention and isolation rules
  • Create and maintain ingestion pipelines: web enrichment, orders, machine telemetry, accounting feeds, idempotent and replayable
  • Ensure data quality with contracts, tests, anomaly detection and stewardship workflows
  • Develop canonical entity model: one identity per person, company, design, part, order and machine
  • Implement deterministic keys, blocking and scoring for entity resolution, review queue and reversible merges
  • Track provenance: source, confidence and history on derived records and corrections that survive reruns
  • Handle time-series and events: taxonomy, downsampling, retention tiers, reliability metrics
  • Develop geospatial features and serving: geocoding, drive-time catchments, facility siting and analytics layer

Skills

Data engineering
Pipelines design
Entity resolution
Information modeling
SQL
Python
Query optimization
System design
Geospatial analysis
Time-series

Education

Bachelor's degree in CS or related field

Tools

SQL
Python
Document stores
Relational stores
Graph databases

Job description

What to Expect

Your role is to build the infrastructure that collects, stores and models everything the company knows about its customers, its machines and itself, and turns it into decisions.

That infrastructure is Cyberdeck: we have rebuilt in our own image every operational software product a company like ours would buy, and run it on our own hardware. Owning the software keeps the data in one estate under one identity model for people and agents.

You will own the stores and the pipelines that fill them, quality and provenance as properties of the pipeline, the entity model beneath it, the time-series and geospatial work the company asks of that data, and the serving layer people and agents read through.

We are looking for data engineers with experience and interest in building and operating production data platforms, information modeling and entity resolution, provenance and data quality, time-series pipelines, and geospatial analysis.

Every role works with our AI systems daily. What can be deterministic, must be. You need no AI background; we prefer people without one. We hire for your knowledge and experience in the field first, so you can steer the ship; the tooling is a learning curve we expect you to take on.

We work on-site in El Segundo. This is hard work, but you will be rewarded with equity in a company we believe will become one of the world's most valuable.

What You'll Do
  • The data estate: stores, schemas and indexes, migrations, query performance, retention and isolation rules
  • Ingestion pipelines: web enrichment, customer orders, machine telemetry, accounting feeds, idempotent, replayable and backfillable
  • Data quality: contracts and tests at the boundary, anomaly detection, freshness ranking, stewardship workflows
  • The canonical entity model: one identity per person, company, design, part, order and machine
  • Entity resolution: deterministic keys, blocking and scoring, a review queue, and reversible merges
  • Provenance: source, confidence and history on derived records, and corrections that survive a rerun
  • Time-series and events: a taxonomy for the floor, downsampling, retention tiers, and reliability metrics
  • Geospatial and serving: geocoding, drive-time catchments, facility siting near customers, an analytics layer
What You'll Bring
  • A degree in computer science, data engineering, information systems or a related field, or equivalent experience
  • Hands-on production data pipelines you were responsible for: idempotency, replay, backfills and schema evolution
  • Information modeling and entity resolution someone else had to live with: canonical identifiers, matching, reversible merges
  • Fluent with document and relational stores, SQL and Python, and able to explain a slow query
  • The judgment to take a wrong number back through the pipeline instead of correcting the output
  • Ready to build the first version of a system and write it for the next engineer
  • Bonus: knowledge graphs or taxonomy design, geospatial work such as isochrones and catchments, or high-rate time-series
Expected Compensation

$100k to $250k base salary, plus equity

  • Equity participation in a high-growth startup
  • Comprehensive health, dental, and vision insurance
  • 401(k) with company matching
  • Bonus for living within 5 miles of our El Segundo facility
  • On-site work with rare work-from-home exceptions
  • Merit-based organization where contribution drives reward
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineer
Software Engineer

Nebula • El Segundo (CA)

Hybrid
USD 100,000 - 250,000
Equity
Health insurance
Dental insurance
+3
Data Engineer
Data Engineer

Baselayer • San Francisco (CA)

Hybrid
USD 120,000 - 150,000
Flexible PTO
SF-based office 4 days/week
Equity
+3
Data Engineer Engineering San Francisco, California
Data Engineer Engineering San Francisco, California

Baselayer • San Francisco (CA), Northern (KY)

Hybrid
USD 120,000 - 150,000
Flexible PTO
SF office
Competitive pay
+2
Infrastructure Engineer
Infrastructure Engineer

Nebula • El Segundo (CA)

Hybrid
USD 100,000 - 250,000
Equity participation
Health, dental, vision
401(k) with matching
+2
Senior Data Analyst
Senior Data Analyst

Bobyard • San Francisco (CA), Northern (KY)

On-site
USD 140,000 - 160,000
Competitive base salary
Equity
Medical, dental, and vision
+1
Data Engineer
Data Engineer

youcom • San Francisco (CA)

Hybrid
USD 180,000 - 220,000
In-person gatherings in SF & NYC
Data Platform Engineer
Data Platform Engineer

Worth AI, Inc. • Orlando (FL), Northern (KY)

Hybrid
USD 120,000 - 180,000
Health Plan
Retirement Plan
Life Insurance
+6
Data Platform Engineer
Data Platform Engineer

Worth AI • Orlando (FL)

Hybrid
USD 120,000 - 180,000
Health Care Plan
Retirement Plan
Life Insurance
+6
Senior Data Engineer
Senior Data Engineer

X4 Engineering • Palo Alto (CA)

Hybrid
USD 150,000 - 220,000
Comprehensive health coverage
Flexible time off
Strong parental support
+2
Founding Data Engineer
Founding Data Engineer

Inventure • San Francisco (CA)

On-site
USD 120,000 - 190,000
Equity