Senior Software Engineer, Data Platform

Matterworks

Somerville (MA)

Hybride

USD 150 000 - 210 000

Plein temps

14 jours+
Générateur de candidature

Obtenez une réponse de cet employeur — un CV et une lettre de motivation adaptés exactement à ce qu’il recherche.

Passez les filtres ATS

Avantages offerts par ce poste

Health insurance
Dental insurance
Vision insurance
Disability insurance
Life insurance
401k with company match
Commuter benefits
Parking
Unlimited time away
Education/conference support

Résumé du poste

Matterworks is seeking a Senior Software Engineer to build the connective tissue of our data platform, designing and scaling systems that enrich data from raw samples into usable datasets with biological context. You will work daily with ML researchers, scientists, and product teams to serve data for research and product development.

You will own data contracts, enable scalable data processing, and drive production-quality checks, cost-aware delivery, and robust interfaces for cross-disciplinary

Qualifications

  • Significant professional experience building production data systems and pipelines.
  • Proficient in Python and SQL for large-scale data processing.
  • Proficient in Kubernetes-native batch orchestration and modern data lake technologies.
  • Demonstrated experience designing stable identifiers for a large, changing corpus.
  • Experience putting an LLM or agent component into a production data path.
  • Daily use of AI coding tools with caution on production data correctness.
  • ownership of work through to running, validated systems.
  • Comfort with messy scientific formats and toolchains.
  • Autonomy and eagerness to learn in an early-stage startup.

Responsabilités

  • Build and Scale Data Contracts: Own the pipelines and system others consume from.
  • Serve and Cost at Scale: Build and scale systems to acquire, store and serve data efficiently.
  • Labels and Enrichment: Turn raw data into usable datasets with consistent schemas and metadata.
  • Quality Gates: Automate quality checks that ensure progress with no regression.
  • Interfaces People Use: Own surfaces called by SDKs and tools exposing platform capabilities.
  • Operations and Data Rights: Ensure SLAs, provenance, and secure data processing.

Connaissances

Python
SQL
Kubernetes
Data pipelines
Large-scale data
Argo Workflows
Metaflow
EKS
Glue
Athena
Iceberg
Parquet
DuckDB
Terraform

Outils

Argo Workflows
Metaflow
EKS
Glue
Athena
Apache Iceberg
Parquet
DuckDB
Terraform
Airflow
Dagster

Description du poste

About Us

Most of the molecules driving human biology are invisible to us. Mass spectrometers already detect metabolites, lipids, and peptides, but the vast majority of those signals never get identified. A typical experiment names a small fraction of its features and discards the rest. We call this biology's dark matter. It’s signal-rich and mechanism-defining, yet almost entirely opaque. Matterworks is building the foundation models that make that dark matter legible. Our Large Spectral Models do for biochemical biology what AlphaFold and ESM did for proteins: turning a library-bound discipline into something predictable and generative, and embedding it at every stage of R&D. Come build the future of biological discovery with us.

Position Overview

As a Senior Software Engineer you’ll work to build the connective tissue of our data platform, in both what you build and how you build it. Design, build and scale systems to enrich data from raw samples and information into readily usable datasets enriched with biological context. The data produced by you and the team will serve our customers through both our ML research and model development activities as well as our product. You will report to the Head of Engineering and work daily with our machine learning researchers, scientists, and product team.

Key Responsibilities
  • Build and Scale Data Contracts: Own the pipelines and system other teams consume from. Design and implement systems that scale to multiple petabytes of data in an effective way.
  • Serving and Cost at Scale: You will build and scale systems to acquire, store and serve data in a fast and affordable as it grows: data layout, featurization throughput, Kubernetes-native orchestration, and cost surfaced before it adds up.
  • Labels and Enrichment: Turn raw data into datasets people can use, with consistent schemas, trustworthy metadata, and documented definitions. Scale scientific labels from studies down to their spectra and underlying features.
  • Quality Gates: Automate quality checks that enable increasing capability without regression and promote only on a pass.
  • Interfaces People Use: Own the surfaces AI, chemistry, product, and agents call, from the SDK used to build datasets to the tools that expose platform capabilities.
  • Operations and Data Rights: Ensure effective operations of our data needs meeting our designed service level agreements, while providing high quality, provenance, and secure data processing in line with our customer needs.
About You
  • Significant professional experience building production data systems and pipelines.
  • We level on scope and judgment rather than years.
  • Proficient in Python and SQL for large-scale data processing.
  • Proficient in Kubernetes-native batch orchestration and modern data lake technologies (Argo Workflows, Metaflow, EKS, Glue, Athena, Apache Iceberg, Parquet, DuckDB, Terraform). Airflow or Dagster experience transfers fine.
  • Demonstrated experience designing stable identifiers for a large, changing corpus, and building validation that gates a publish rather than reporting on it after the fact.
  • Experience putting an LLM or agent component into a production data path, including the eval loop, the gold set, and cost per record.
  • Daily use of AI coding tools, paired with healthy skepticism about their output on questions of production data correctness.
  • A track record of owning work through to a running, validated system, including the unglamorous parts: fixing the malformed dataset, writing the backfill, debugging last night's bad publish.
  • Comfort with messy scientific formats and toolchains (mzML, RDKit, ProteoWizard or similar). Engineering depth is the requirement.
  • A passion for contributing to an early-stage startup where autonomy, eagerness to learn, and enthusiasm for solving novel scientific challenges prevail over rigid processes and egos.
Working at Matterworks

Given the cross-disciplinary and innovative nature of our work, effective collaboration and communication are critical to our progress. We operate in a flexible hybrid model that accommodates both fully remote team members and those who work full-time from our Somerville, MA office. While some positions may require regular in-person presence for hands-on work or local collaboration, many roles can be performed remotely with team members distributed across various locations.

Compensation and Benefits
  • Matterworks offers full-time employees a competitive base salary, stock options, and benefits (health & dental, vision, long- and short-term disability, life insurance, 401k with company match).
  • Employees enjoy a flexible work & unlimited time away policy, commuter benefits and parking, regular team meals and outings, and company support for continued education/coursework and conference participation.

Matterworks, Inc. is an equal opportunity employer.

All candidates for employment at Matterworks are considered without regard to race, color, religion, national origin, age, sex, marital status, ancestry, physical or mental disability, veteran status, gender identity, sexual orientation, or any other category protected by law.

Obtenez votre examen gratuit et confidentiel de votre CV.

ou faites glisser et déposez votre fichier ici.

Similar jobs

Postes similaires à comparer

Data Engineer
Data Engineer

Matterworks • Somerville (MA)

Sur place
USD 120 000 - 170 000
Health insurance
Stock options
401k with company match
+4
Full-Stack Software Engineer
Full-Stack Software Engineer

Matter • El Segundo (CA)

Hybride
USD 120 000 - 170 000
Competitive compensation
Early-stage equity
Health, dental, vision coverage
Software Engineer, Data Platform
Software Engineer, Data Platform

General Matter • Los Angeles (CA)

Sur place
USD 125 000 - 220 000
Stock options
Medical, vision & dental coverage
401(k) retirement plan
Founding Computational Scientist, Proteomics Onsite (Cambridge, MA)
Founding Computational Scientist, Proteomics Onsite (Cambridge, MA)

S27a • Cambridge (MA)

Sur place
USD 120 000 - 180 000
Member of the Technical Staff, Biological Data
Member of the Technical Staff, Biological Data

Output Biosciences • San Francisco (CA)

Sur place
USD 160 000 - 230 000
Equity stake
Comprehensive benefits
Ownership culture
Head of Marketing
Head of Marketing

Matter Intelligence • San Francisco (CA)

Sur place
USD 180 000 - 260 000
Equity
Health coverage (100% employer-paid)
Unique opportunity to work on novel AI
+1
Member of the Technical Staff, Biological Data
Member of the Technical Staff, Biological Data

Output Biosciences • New York (NY)

Sur place
USD 80 000 - 120 000
Competitive salary and equity
Excellent medical, dental, and vision coverage
AI Engineer / Senior AI Engineer
AI Engineer / Senior AI Engineer

Blue Matter • South San Francisco (CA)

Sur place
USD 120 000 - 170 000
401k with company match
Medical, dental, and vision
FSA/HSA
+3
Applied AI Engineer (Product)
Applied AI Engineer (Product)

Matter Intelligence • San Francisco (CA)

Sur place
USD 140 000 - 210 000
Equity
Health coverage
Onsite work allowance
Senior Data Engineer II, Product Engineering
Senior Data Engineer II, Product Engineering

Polygon.io, Inc • Northern (KY)

Sur place
USD 120 000 - 180 000
401(k) plan
Unlimited time off
Medical plans