Data Engineer

Function

United States

On-site

USD 120,000 - 170,000

Full time

5 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

401k Match
Health Insurance
Choose your own holidays

Job summary

Function is building a platform that enables subject-matter experts to safely change data, with guardrails, CI for data, and clear lineage. Our three surfaces—tracking infrastructure, data processing, and ML infrastructure—enable fast, safe iteration for engineers and analysts.

We emphasize self-service, templates, and policy-as-code for PHI and data governance. We value strong Python/SQL, experience with Databricks, and familiarity with data contracts, streaming, and MLOps.

Qualifications

  • 1-4 years of engineering experience
  • Strong Python and SQL development skills
  • Experience with data platforms and data pipelines
  • Familiarity with data contracts and lineage tooling

Responsibilities

  • Design interfaces and schemas used by multiple teams
  • Build and maintain internal platform infrastructure
  • Focus on testing and CI for data/ML systems
  • Ensure safe, auditable data changes across environments
  • Contribute to self-service tooling for engineers and analysts

Skills

Python
SQL
Databricks
Data contracts
Data diffing
Lineage tooling
MLOps
Terraform
Dagster
Airflow
Kafka
Spark

Tools

Databricks
Snowflake
BigQuery
dbt
DLT
Dagster
Airflow
Kafka

Job description

  • At Function, the person who best understands a piece of data is rarely the person who can safely change it
  • A data scientist knows exactly how a biomarker should be derived
  • A clinician knows which reference range is wrong
  • Today, both have to file a ticket and wait for an engineer
  • We think that’s a platform failure, not a process problem
  • Our job is to build the paved road that lets the subject matter expert make the change themselves – and lets an agent make it too – without anyone lying awake wondering about the validity or stability of what just shipped
  • That means the guardrails have to be real: declarative contracts instead of hand-rolled code, CI that diffs the data and not just the diff, validations that gate promotion, lineage that tells you who’s downstream, and a rollback that takes one command
  • Get that right and both your humans and your agents get faster at the same time, for the same reason. That’s the work. It’s platform engineering, and the users are engineers
  • Three surfaces, all with the same goal – make the safe path the fast path:
  • Tracking infrastructure. The event pipeline behind product analytics, experimentation, and feature gates. Schemas that are enforced at the source, so a bad event never becomes a bad metric
  • Data processing infrastructure. The Bronze Silver Gold layer in Databricks. Automated schema evolution, contract tests, backfills that aren’t scary, freshness and volume monitors generated from the contract rather than bolted on after
  • ML infrastructure. Feature computation and serving, training and eval pipelines, and the plumbing that gets model output back into the product — with the same testing and observability bar as everything else
  • Cutting across all three. The self-service story. Templates, local dev and preview environments, policy-as-code for PHI, ownership routing for alerts, and progressive gates so an exploratory model ships freely while a member-facing one earns more scrutiny
Benefits
  • 401k Match
  • Health Insurance
  • Choose your own holidays
  • Designed interfaces and schemas that other teams depend on, then evolved them without breaking those teams
  • Built internal platform or infrastructure that other engineers actually adopted. You’ve felt the difference between shipping a tool and getting it used
  • Thought hard about testing and CI for data or ML, where correctness is statistical and the failure is often silentOperated production data or ML systems, on call for them, and fixed them under pressure
  • Strong Python and SQL. Comfortable in a lakehouse – we use Databricks; Snowflake or BigQuery translates fine
  • 1-4 years of engineering experience gets you here, but we care about what you’ve built, not the number
  • Declarative pipeline frameworks (dbt, DLT, Dagster, Airflow)
  • Streaming (Kafka, Spark Structured Streaming)
  • Terraform
  • Data contracts, data diffing, or lineage tooling
  • Feature stores
  • MLOps and eval tooling
  • Agentic coding workflows
  • Healthcare PHI, or HIPAA experience
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Staff Software Engineer - Backend
Staff Software Engineer - Backend

Databricks • Mountain View (CA)

On-site
USD 192,000 - 260,000
Comprehensive benefits
Annual performance bonus
Equity opportunities
Data Engineer
Data Engineer

InfoVision Inc. • Detroit (MI)

On-site
USD 105,000 - 155,000
Staff Software Engineer - Backend
Staff Software Engineer - Backend

Menlo Ventures • Mountain View (CA)

On-site
USD 192,000 - 260,000
Annual performance bonus
Equity options
Comprehensive benefits package
Senior Software Engineer (Data Platform)
Senior Software Engineer (Data Platform)

globalizationpartners • United States

Remote
USD 120,000 - 170,000
Health insurance
Parental leave
Sabbatical after 5 years
Senior Software Engineer - Distributed Data Systems
Senior Software Engineer - Distributed Data Systems

Databricks • Bellevue (WA)

On-site
USD 157,700 - 213,800
Comprehensive benefits
Equity opportunities
Annual performance bonuses
Sr. Software Engineer - Ingestion Core team
Sr. Software Engineer - Ingestion Core team

Cacheflow • San Francisco (CA)

On-site
USD 166,000 - 225,000
Sr Software Engineer, Infrastructure
Sr Software Engineer, Infrastructure

Databricks • San Francisco (CA)

On-site
USD 136,300 - 187,450
Staff Software Engineer - Distributed Data Systems
Staff Software Engineer - Distributed Data Systems

Databricks • Bellevue (WA)

On-site
USD 182,400 - 247,000
Comprehensive benefits
Performance bonuses
Equity options
Senior Software Engineer - Distributed Data Systems
Senior Software Engineer - Distributed Data Systems

Menlo Ventures • Bellevue (WA)

On-site
USD 157,700 - 213,800
Competitive compensation
Annual performance bonus
Equity options
+1
Staff Fullstack Engineer, Agentic Applications
Staff Fullstack Engineer, Agentic Applications

Databricks • Mountain View (CA)

On-site
USD 192,000 - 260,000
Comprehensive benefits
Equity opportunities
Annual performance bonus