Member of Technical Staff, Data Engineering

Mithrl

San Francisco (CA)

On-site

USD 180,000 - 220,000

Full time

4 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Health insurance
401(k)

Job summary

Mithrl in San Francisco seeks a data engineer to own the ingestion and normalization layer for diverse biological data.

You will build robust pipelines, combine classical engineering with LLM-powered cleanup, and collaborate with product, bioinformatics, and infra teams to ensure clean, analysis-ready datasets.

The role offers a mission-driven environment, high ownership, and a chance to impact life sciences through scalable data tooling.

Qualifications

  • 5+ years of experience in data engineering / data wrangling with real-world tabular or semi-structured data.
  • Strong Python proficiency and data-processing libraries (Pandas, Polars, PyArrow).
  • Hands-on experience cleaning messy Excel/CSV data and standardizing it.
  • Experience designing robust ETL/ELT pipelines for scientific or lab data.
  • Ability to combine traditional data engineering with LLM-powered data normalization.
  • Experience owning ingestion and normalization end-to-end with maintainability and scalability.

Responsibilities

  • Build and own an AI-powered ingestion pipeline to import data from various sources.
  • Develop schema mapping, coercion, and unit normalization during ingestion.
  • Structure semi-structured data with metadata extraction and cleaning.
  • Ensure core transformations run at ingestion for clean downstream data.
  • Build validation and QA layers to catch corrupt data before entry.
  • Collaborate with product, data science, and infra to enforce data standards.

Skills

5+ years data engineering
Python
Data wrangling

Tools

Pandas
Polars
PyArrow

Job description

About Mithrl

We envision a world where novel drugs and therapies reach patients in months, not years, accelerating breakthroughs that save lives.

Mithrl is building the world’s first commercially available AI Co-Scientist—a discovery engine that empowers life science teams to go from messy biological data to novel insights in minutes. Scientists ask questions in natural language, and Mithrl answers with real analysis, novel targets, and patent-ready reports.

About Mithrl

We envision a world where novel drugs and therapies reach patients in months, not years, accelerating breakthroughs that save lives.

Mithrl is building the world’s first commercially available AI Co-Scientist—a discovery engine that empowers life science teams to go from messy biological data to novel insights in minutes. Scientists ask questions in natural language, and Mithrl answers with real analysis, novel targets, and patent-ready reports.

Our traction speaks for itself:
  • 12X year-over-year revenue growth
  • Trusted by leading biotechs and big pharma across three continents
  • Driving real breakthroughs from target discovery to patient outcomes.
What You Will Do

Build and own an AI-powered ingestion & normalization pipeline to import data from a wide variety of sources - unprocessed Excel/CSV uploads, lab and instrument exports, as well as processed data from internal pipelines. Develop robust schema mapping, coercion, and conversion logic (think: units normalization, metadata standardization, variable-name harmonization, vendor-instrument quirks, plate-reader formats, reference-genome or annotation updates, batch-effect correction, etc.). Use LLM-driven and classical data-engineering tools to structure "semi-structured" or messy tabular data - extracting metadata, inferring column roles/types, cleaning free-text headers, fixing inconsistencies, and preparing final clean datasets. Ensure all transformations that should only happen once (normalization, coercion, batch-correction) execute during ingestion - so downstream analytics / the AI "Co-Scientist" always works with clean, canonical data. Build validation, verification, and quality-control layers to catch ambiguous, inconsistent, or corrupt data before it enters the platform. Collaborate with product teams, data science / bioinformatics colleagues, and infrastructure engineers to define and enforce data standards, and ensure pipeline outputs integrate cleanly into downstream analysis and storage systems.

What You Bring

Must-have

  • 5+ years of experience in data engineering / data wrangling with real-world tabular or semi-structured data.
  • Strong fluency in Python, and data processing tools (Pandas, Polars, PyArrow, or similar).
  • Excellent experience dealing with messy Excel / CSV / spreadsheet-style data - inconsistent headers, multiple sheets, mixed formats, free-text fields - and normalizing it into clean structures.
  • Comfort designing and maintaining robust ETL/ELT pipelines, ideally for scientific or lab-derived data.
  • Ability to combine classical data engineering with LLM-powered data normalization / metadata extraction / cleaning.
  • Strong desire and ability to own the ingestion & normalization layer end-to-end - from raw upload -> final clean dataset - with an eye for maintainability, reproducibility, and scalability.
  • Good communication skills; able to collaborate across teams (product, bioinformatics, infra) and translate real-world messy data problems into robust engineering solutions.

Nice-to-have

  • Familiarity with scientific data types and "modalities" (e.g. plate-readers, genomics metadata, time-series, batch-info, instrumentation outputs).
  • Experience with workflow orchestration tools (e.g. Nextflow, Prefect, Airflow, Dagster), or building pipeline abstractions.
  • Experience with cloud infrastructure and data storage (AWS S3, data lakes/warehouses, database schemas) to support multi-tenant ingestion.
  • Past exposure to LLM-based data transformation or cleansing agents - building or integrating tools that clean or structure messy data automatically.
  • Any background in computational biology / lab-data / bioinformatics is a bonus - though not required.
What You Will Love At Mithrl
  • Mission-driven impact: you’ll be the gatekeeper of data quality - ensuring that all scientific data entering Mithrl becomes clean, consistent, and analysis-ready. You’ll have outsized influence over the reliability and trustworthiness of our entire data + AI stack.
  • High ownership & autonomy: this role is yours to shape. You decide how ingestion works, define the standards, build the pipelines. You’ll work closely with our product, data science, and infrastructure teams - shaping how data is ingested, stored, and exposed to end users or AI agents.
  • Team: Join a tight-knit, talent-dense team of engineers, scientists, and builders
  • Culture: We value consistency, clarity, and hard work. We solve hard problems through focused daily execution
  • Speed: We ship fast (2x/week) and improve continuously based on real user feedback
  • Location: Beautiful SF office with a high-energy, in-person culture
  • Benefits: Comprehensive PPO health coverage through Anthem (medical, dental, and vision) + 401(k) with top-tier plans

Not all strong candidates will meet every single qualification as listed.

Research shows that people who identify as being from underrepresented groups are more prone to experiencing imposter syndrome and doubting the strength of their candidacy. We think AI systems like the ones we're building have enormous social and ethical implications. We think this makes representation even more important, and we strive to include a range of diverse perspectives on our team.

Compensation Range: $180K - $220K

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Data Engineer, Scientific Data Ingestion
Data Engineer, Scientific Data Ingestion

Mithrl • San Francisco (CA)

On-site
USD 120,000 - 150,000
Comprehensive PPO health coverage
401(k) with top-tier plans
High ownership and autonomy
Member of Technical Staff, Discovery Applications
Member of Technical Staff, Discovery Applications

Mithrl • San Francisco (CA)

On-site
USD 140,000 - 210,000
Health insurance
401(k)
Platform Solutions Engineer
Platform Solutions Engineer

Mithrl • San Francisco (CA)

On-site
USD 180,000 - 220,000
Health coverage through Anthem
401(k) with top-tier plans
On-site SF office
Member of Technical Staff, Biological Analysis & Simulation
Member of Technical Staff, Biological Analysis & Simulation

Mithrl • San Francisco (CA)

On-site
USD 150,000 - 190,000
Health coverage
401(k) plan
Enterprise Account Executive
Enterprise Account Executive

Mithrl • San Francisco (CA), Northern (KY)

Hybrid
USD 140,000 - 210,000
Comprehensive health coverage
Senior AI/ML Engineer
Senior AI/ML Engineer

Prudentia Sciences • Boston (MA)

On-site
USD 206,000 - 240,000
Remote-friendly
Competitive compensation
Equity
Analytics Engineer
Analytics Engineer

Thesis • New York (NY)

On-site
USD 120,000 - 150,000
Equity package
Health insurance
HSA and pre-tax benefits
+6
Lead, Data AI & Engineering
Lead, Data AI & Engineering

MHC Automation • Burnsville (MN), Northern (KY)

Hybrid
USD 170,000 - 230,000
Workplace Flexibility
401(k) Plan with employer match
Medical, Dental, Vision plans
+2
Member of Technical Staff - Forward Deployed Scientist
Member of Technical Staff - Forward Deployed Scientist

Phylo • South San Francisco (CA)

On-site
USD 170,000 - 275,000
Competitive salary
Full medical, dental, and vision
401(k)
+4
Senior Member of Technical Staff (Applied AI)
Senior Member of Technical Staff (Applied AI)

Solstice • New York (NY)

On-site
USD 230,000 - 350,000
Health, dental, vision insurance
Equity opportunity
Visa sponsorship (O-1, H-1B, TN)
+5