Software Engineer [ Data Pipelines & Interface ]

Metamorphic

Palo Alto (CA)

On-site

USD 160,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Visa sponsorship
Competitive compensation
Mentorship and career development

Job summary

Metamorphic is seeking Software Engineers to build a robust data platform powering research and model development. You will design and maintain systems moving data from ingestion through transformation to analytics-ready interfaces for research workflows, model training, and long‑term archival use.

You will collaborate with researchers, ML engineers, and infrastructure teams to ensure production readiness, scalability, and maintainability as the platform evolves with the company.

Qualifications

  • Bachelor’s degree or equivalent in CS/ML or related field.
  • Strong Python and database query skills.
  • Experience with production data platforms and ETL/ELT pipelines.
  • Experience with relational and NoSQL databases.

Responsibilities

  • Design and maintain the data platform for analytics and research workflows.
  • Build and optimize ETL/ELT pipelines and ingestion services.
  • Ensure data quality, schema evolution, and performance across SQL/NoSQL.
  • Collaborate with researchers, ML engineers, and infra teams.

Skills

Python
Data pipelines
SQL/NoSQL
Data modeling
Observability
Collaboration
Airflow
Dagster
Prefect
Performance optimization

Education

Bachelor's degree in Computer Science, ML, or related field

Tools

Airflow
Dagster
Prefect
Docker
Kubernetes
MLflow

Job description

About The Role

We are hiring Software Engineers to build the data platform that powers our research and model‑development efforts. You will help design and maintain the systems that move data from acquisition through transformation and storage into reliable downstream interfaces for analytics, research workflows, model training, and long‑term archival use. This includes database design and administration, ETL/ELT pipelines, ingestion services, data modeling, schema evolution, observability, and performance optimization across both SQL and NoSQL systems. This role sits at the intersection of software engineering, data engineering, and scientific infrastructure. You will work closely with researchers, ML engineers, and infrastructure teams to ensure that our data systems are robust enough for production use, flexible enough for evolving research needs, and maintainable enough to serve as a foundation for years of scientific and engineering work. You will have substantial autonomy in shaping how our long‑lived scientific data platform is represented, administered, and evolved as the company scales.

You’ll Thrive In This Role If You
  • Are excited about working in a fast‑paced, production‑focused research lab that often requires switching between many hats
  • Have significant software engineering experience and can move quickly without sacrificing rigor
  • Are able to balance research goals with practical engineering constraints
  • Enjoy pair programming and deeply collaborative work
  • Are eager to learn more about machine learning research in a novel scientific domain
  • Are enthusiastic to work at an organization that functions as a single, cohesive team pursuing large‑scale AI research
  • Have ambitious goals for AI progress and are excited to create the best outcomes over the long term
We Offer
  • The chance to work on one of the most scientifically consequential AI projects being pursued today
  • A small, world‑class team where your contributions directly shape the science and the company
  • Competitive compensation and benefits, along with visa sponsorship
  • Strong mentorship and career development
Salary Range

$160,000 - $240,000 USD
Based on experience. We additionally offer a competitive equity package and comprehensive benefits, as well as visa sponsorship for international candidates.

Minimum Qualifications
  • Bachelor’s degree or equivalent experience in Computer Science, Machine Learning, Computational Neuroscience, or a related field
  • Strong software engineering skills and strong working proficiency in Python and database query languages
  • Experience designing, building, and maintaining production object stores and ETL/ELT pipelines
  • Deep familiarity with relational and NoSQL database systems, including schema design, indexing, query optimization, and administration
  • Experience designing durable data models and schemas for complex, evolving datasets
  • Experience maintaining data platforms in production environments, including tools for monitoring, observability, troubleshooting, and incident response
  • Experience building high‑throughput ingestion systems and optimizing data movement across storage, compute, and network boundaries
  • Familiarity with workflow orchestration tools (e.g. Airflow, Dagster, Prefect)
  • Experience collaborating closely with various teams simultaneously (e.g. research, ML, scientific, etc.) to translate ambiguous requirements into robust, consensus‑driven data pipeline specs
Nice to Have
  • Experience serving as a database administrator (DBA) or having substantial DBA‑style ownership of production data systems
  • Experience with scientific, biomedical, behavioral, or neural datasets, especially where data provenance and long‑term reuse matter
  • Experience with performance‑critical compiled or systems languages (e.g. Rust, Zig, C++)
  • Experience with containerization, and scaling container orchestration (e.g. via Docker, Kubernetes)
  • Experience with provisioning and managing distributed GPU compute infrastructure (e.g. via Ray, Dstack, Skypilot)
  • Proficiency with MLOps platforms for experiment tracking and reproducibility (e.g. MLflow, W&B)
  • Background in scientific computing, computational neuroscience, life sciences, or ML‑adjacent research environments
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Research Engineer [ Data Engineering ]
Research Engineer [ Data Engineering ]

Metamorphic • Palo Alto (CA)

On-site
USD 175,000 - 250,000
Software Engineer, Data Infrastructure
Software Engineer, Data Infrastructure

Thinkingmachines • San Francisco (CA)

On-site
USD 350,000 - 475,000
Health, dental, and vision insurance
Unlimited paid time off (PTO)
Paid parental leave
+1
Engineering Manager, Research Data Platform
Engineering Manager, Research Data Platform

United States Digital Space LLC • San Francisco (CA)

Hybrid
USD 405,000 - 850,000
Data Engineer
Data Engineer

Southern Arkansas University • Warner Robins (GA)

Remote
Flexible hours
Weekly bonus of $500–$1000 USD
Work from anywhere
Software Engineer, Research Infrastructure
Software Engineer, Research Infrastructure

United States Digital Space LLC • San Francisco (CA)

Hybrid
USD 405,000 - 625,000
Principal Software Engineer, Data
Principal Software Engineer, Data

aijoblist • San Francisco (CA), Cambridge (MA)

On-site
USD 204,000 - 348,000
Medical, dental, and vision coverage
Flexible time off
Paid parental leave
+2
Staff Software Engineer
Staff Software Engineer

MilliporeSigma • St. Louis (MO)

On-site
USD 110,000 - 166,000
Health insurance
Paid time off (PTO)
Retirement contributions
Software Engineer, Data/ML
Software Engineer, Data/ML

XOXO AI Inc. • San Francisco (CA)

On-site
USD 250,000 - 500,000
Top-tier health benefits
Dental & vision coverage
Equity 1–5%
Senior Data Engineer
Senior Data Engineer

Intercontinental Exchange (ICE) • Atlanta (GA)

On-site
USD 100,000 - 140,000
Senior Software Engineer, Data
Senior Software Engineer, Data

Lila Sciences • San Francisco (CA), Cambridge (MA)

On-site
USD 144,000 - 288,000
Comprehensive benefits program
Flexible time off
Paid parental leave
+2