Data Engineer, Foundations

Superhuman Labs, Inc.

San Francisco (CA)

Hybrid

USD 175,000 - 245,000

Full time

6 hours ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Hybrid work model
Health, dental, vision benefits
401(k) matching
Paid parental leave

Job summary

Superhuman Labs, Inc. seeks a Data Engineer to join the Data Foundations team, building robust, scalable data pipelines and data lakes to process billions of daily events. You will shape architecture and contribute to production-grade data infrastructure across product surfaces.

The role blends data engineering, data-intensive applications, and data modeling, with ownership from strategy to production. Hybrid work is offered with preferred SF or Seattle locations.

Qualifications

  • 3+ years of experience running live production environments, including high-load or data-intensive workflows, with a focus on uptime and reliability.
  • Proficient in SQL and Python, with deep hands-on experience in Spark and a modern lakehouse or cloud data warehouse (Databricks, Delta Lake, dbt, Snowflake, or similar).
  • Strong knowledge of ETL/ELT design patterns, orchestration tools (Airflow, dbt, or Dagster), data quality frameworks, and CI/CD for data with Git-based deployments.
  • Data modeling and data warehouse design skills, along with a rigorous approach to data quality and observability.
  • Hands-on experience with modern data storage technologies (Delta Lake, Snowflake, BigQuery, or Redshift).
  • Clear communication and collaboration across diverse teams turning business needs into robust data solutions and trustworthy metrics.
  • Strong ownership mindset, end-to-end responsibility for systems from strategy to production and reliability.
  • Strong acumen for scalable, complex data processing and prioritization across multiple projects in a fast-paced environment.

Responsibilities

  • Design and implement robust, scalable data pipelines and systems for large data volumes.
  • Build large-scale data pipelines and data lakes (Spark/Databricks) ingesting billions of daily events.
  • Design solutions to keep data available, secure, and scalable for real-time and batch processing.
  • Own data quality, freshness, and reliability for foundational datasets with automated checks and alerts.
  • Build ETL frameworks and tooling for other data teams and data scientists to self-serve datasets.
  • Collaborate with product teams, back-end engineers, ML engineers, and data science to turn business questions into data products.
  • Partner with leadership to shape the team’s charter and technical roadmap.

Skills

SQL
Python
Spark
Databricks
Delta Lake
dbt
Snowflake
Airflow
Dagster
Git

Tools

Airflow
dbt
Dagster
Git

Job description

Engineering, Product, Design, and Marketing Engineering Data

Superhuman offers a dynamic hybrid working model for this role. This flexible approach gives team members the best of both worlds: plenty of focus time along with in-person collaboration that helps foster trust, innovation, and a strong team culture.

Preferred locations for team members for our Data Platform roles are our San Francisco or Seattle hubs.

About Superhuman

Grammarly is now part of Superhuman, the AI productivity platform on a mission to unlock the superhuman potential in everyone. The Superhuman suite of apps and agents brings AI wherever people work, integrating with over 1 million applications and websites. The company's products include Grammarly's writing assistance, Docs' collaborative workspace, Mail's inbox management, and Go, the proactive AI assistant that understands context and delivers help automatically. Founded in 2009, Superhuman empowers over 40 million people, 50,000 organizations, and 3,000 educational institutions worldwide to eliminate busywork and focus on what matters. Learn more at superhuman.com and about our values here .

The Opportunity

To achieve our ambitious goals, we’re looking for a Data Engineer to join our Data Foundations team and help us build a world-class data platform. Superhuman's success depends on its ability to efficiently ingest and process over 70 billion daily events to improve our products. This role presents a unique opportunity to experience all aspects of building complex data pipelines and systems, including contributing to the strategy, defining the architecture, and developing and shipping to production.

Superhuman is a compound startup: we build many products as one integrated suite rather than standalone tools. That model creates an unusually rich data opportunity, with signals spanning our full product suite across both consumers and enterprises - and the foundational data you build connects those surfaces.

The Data Foundations team is part of the Data Platform org, powering data needs across Superhuman rather than a single product surface. It owns the foundational datasets and data models that define the fundamental entities of Superhuman's business, serving as building blocks for the company's other data teams and data scientists. Behind the scenes, the team develops and runs large-scale ETL pipelines that process petabytes of data and billions of events daily.

This is a high-impact role at the intersection of data engineering, data-intensive applications, and data modeling, where the systems you build directly shape how efficiently our other Data teams can build their datasets.

What you’ll do

As a Data Engineer on the Data Foundations team, you will design and implement robust, scalable, and reliable data pipelines and systems that handle large volumes of data, empowering both product features and data-driven decision-making across the company. Your work will span areas such as ETL data pipelines, data lakes, performance-oriented data processing, and ETL framework development.

Architect, build, and own large-scale data pipelines and data lakes (Spark/Databricks) that ingest and process billions of daily events into reliable, decision-grade datasets.

Design and implement solutions that keep data available, secure, and scalable across the platform, enabling both real-time and batch processing.

Own data quality, freshness, and reliability for the foundational datasets other teams depend on, with automated checks, monitoring, and alerting.

Build the ETL frameworks and tooling that let other Data teams and data scientists self-serve and model core business entities into clean, well-documented, reusable datasets.

Partner with product teams, back-end engineers, ML engineers, and Data Science to turn business questions into high-impact data products behind business-critical features, research, and experimentation.

Collaborate with leadership to shape the team’s charter and technical roadmap.

Qualifications

You have 3+ years of experience running live production environments, including high-load or data-intensive workflows, with a focus on uptime and reliability.

You’re proficient in SQL and Python, with deep hands-on experience in Spark and a modern lakehouse or cloud data warehouse (Databricks, Delta Lake, dbt, Snowflake, or similar).

You have strong knowledge of ETL/ELT design patterns, orchestration tools (e.g., Airflow, dbt, or Dagster), data quality frameworks, and CI/CD for data with Git-based deployments.

You have data modeling and data warehouse design skills, along with a rigorous approach to data quality and observability.

You have hands-on experience with modern data storage technologies (for example, Delta Lake, Snowflake, BigQuery, or Redshift).

You communicate clearly and collaborate well across diverse teams and stakeholders, turning business needs into robust data solutions and trustworthy metrics.

You bring a strong ownership mindset, taking end-to-end responsibility for the systems you build, from strategy and architecture through production and ongoing reliability.

You have a strong acumen for scalable, highly complex data processing, think holistically about problems, and raise the bar on your team’s craft.

You’re a self-starting problem-solver who thinks from first principles, manages priorities across multiple projects, and thrives in a fast-paced, results-driven environment.

Nice to have

You have experience operating large-scale distributed systems, as well as in data engineering and data modeling.

Experience with high-throughput, real-time streaming systems (for example, Kafka, Flink, or Spark Structured Streaming) at the scale of billions of events per day.

Comfort with data lake and lakehouse technologies (for example, Delta Lake, Iceberg, or Hudi) and managing cloud infrastructure as code (for example, Terraform).

A track record of building self-serve data products, tools, or frameworks that other teams rely on.

Experience partnering with analytics, data science, or machine learning teams as a strategic data partner to productionize data and models.

Compensation and Benefits

Superhuman offers all team members competitive pay along with a benefits package encompassing the following and more:

Excellent health care (including a wide range of medical, dental, vision, mental health, and fertility benefits)

Disability and life insurance options

401(k) matching

Paid parental leave

20 days of paid time off per year, 12 days of paid holidays per year, two floating holidays per year, and flexible sick time

Generous stipends (including those for caregiving, pet care, wellness, your home office, and more)

Annual professional development budget and opportunities

Superhuman takes a market-based approach to compensation, which means base pay may vary depending on your location. Our US locations are categorized into two compensation zones based on proximity to our hub locations.

Base pay may vary considerably depending on job-related knowledge, skills, and experience. The expected salary ranges for this position are outlined below by compensation zone and may be modified in the future.

US Zone 1

175,000 to 245,000

US Zone 2

157,000 to 220,500

We encourage you to apply

At Superhuman, we value our differences, and we encourage all to apply- especially those whose identities are traditionally underrepresented in tech organizations. We do not discriminate on the basis of race, religion, color, gender expression or identity, sexual orientation, ancestry, national origin, citizenship, age, marital status, veteran status, disability status, political belief, or any other characteristic protected by law. Superhuman is an equal opportunity employer and a participant in the US federal E-Verify program (US). We also abide by the Employment Equity Act (Canada).

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Engineer, Foundations
Data Engineer, Foundations

Zoomcar • San Francisco (CA)

Hybrid
USD 157,000 - 245,000
Health care
Disability and life insurance options
401(k) matching
+8
Data Engineer, Foundations
Data Engineer, Foundations

Superhuman • San Francisco (CA)

On-site
USD 175,000 - 245,000
Health benefits
Life insurance
401(k) matching
+7
Software Engineer, Data Platform
Software Engineer, Data Platform

Grammarly • San Francisco (CA)

On-site
USD 225,000 - 275,000
Health care benefits
401(k) matching
Professional development budget
+2
Software Engineer, Data Engineering
Software Engineer, Data Engineering

Superhuman • San Francisco (CA)

Hybrid
USD 225,000 - 275,000
Excellent health care
Disability and life insurance
401(k) and RRSP matching
+4
Software Engineer, Data Governance
Software Engineer, Data Governance

Superhuman • Seattle (WA), San Francisco (CA)

On-site
USD 180,000 - 260,000
Health benefits
401(k) matching
Paid parental leave
+1
Software Engineer, Data Governance
Software Engineer, Data Governance

Coda • San Francisco (CA)

Hybrid
USD 190,000 - 230,000
Health care
Disability insurance
401(k) matching
+5
Software Engineer, Data Governance
Software Engineer, Data Governance

Superhuman Labs, Inc. • San Francisco (CA)

Hybrid
USD 220,000 - 290,000
Excellent health care
401(k) matching
Paid parental leave
+4
Senior Product Manager, Data Platform
Senior Product Manager, Data Platform

Coda • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Health benefits
Paid time off
Professional development budget
+1
Software Engineer, Superhuman Database Infrastructure
Software Engineer, Superhuman Database Infrastructure

Zoomcar • Seattle (WA)

Hybrid
USD 170,000 - 230,000
Health insurance
Paid time off
Professional development budget
Software Engineer, Superhuman Database Infrastructure
Software Engineer, Superhuman Database Infrastructure

Superhuman Labs, Inc. • Seattle (WA)

Hybrid
USD 164,000 - 284,000
Health care
401(k) matching
Paid parental leave
+1