Senior Software Engineer, Data Platform

Profluent Bio

Emeryville (CA)

On-site

USD 170,000 - 220,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive compensation package
401(k) with employer match
Comprehensive health/dental/vision insurance
Generous PTO policy
Professional development opportunities

Job summary

Profluent Bio is seeking a Senior Software Engineer to design, build, and scale its data platform in Emeryville, CA. This role will focus on creating robust data systems for protein engineering campaigns.

The ideal candidate will have over 5 years of experience and a strong background in Python, cloud platforms, and data security. The position offers competitive compensation, including equity and a strong benefits package.

Qualifications

  • 5+ years of software engineering or data platform experience.
  • Strong proficiency in modern software development practices.
  • Experience designing production data pipelines and data warehouses.

Responsibilities

  • Design, build, and maintain scalable data infrastructure.
  • Develop secure data pipelines with strong access control.
  • Own core components of the data warehouse.

Skills

Python
Data engineering
Cloud platforms (GCP)
Data security
CI/CD

Education

BS, MS, or PhD in Computer Science or related field

Tools

PostgreSQL
BigQuery

Job description

Profluent is an AI-first protein design company. Founded in 2022, we develop deep generative models to design and validate novel, functional proteins to revolutionize biomedicine. Based in Emeryville, CA, we are backed by leading investors including Altimeter Capital, Bezos Expeditions, Spark Capital, Insight Partners, Air Street Capital, AIX Ventures, and Convergent Ventures, and have raised over $150M to date.

We’re looking for a Senior Software Engineer to help design, build, and scale Profluent’s data platform. This platform houses data from protein engineering campaigns, including protein designs, experimental results, partner datasets, analytical outputs, and model‑ready training data. It enables rapid machine learning, biological discovery, and secure collaboration across internal and external programs.

This role is ideal for an engineer who enjoys building robust data systems: secure ingestion pipelines, well‑structured warehouses, reliable data models, access controls, auditability, and infrastructure that makes complex scientific data usable at scale. You will work closely with ML, bioinformatics, and program teams to ensure Profluent’s data is organized, governed, accessible, and protected.

Responsibilities
  • Design, build, and maintain scalable data infrastructure for protein engineering campaigns, including ingestion, transformation, validation, storage, and retrieval of large scientific datasets
  • Develop secure data pipelines for internal and partner-generated data, with strong attention to access control, data siloing, provenance, auditability, and compliance with data use restrictions
  • Own core components of Profluent’s data warehouse and data platform, using Python, GCP, PostgreSQL, BigQuery, and related cloud-native technologies
  • Build systems that transform raw experimental, computational, and partner data into structured, reliable, analysis-ready and model-ready datasets
  • Establish best practices for data modeling, metadata management, data quality checks, schema evolution, versioning, and documentation
  • Collaborate with ML engineers, computational biologists, data scientists, and program stakeholders to understand data requirements and translate them into scalable technical systems
  • Improve engineering quality through thoughtful system design, code review, testing, CI/CD, observability, and maintainable development workflows
  • Contribute to architectural decisions for how Profluent stores, secures, organizes, and uses data across programs and partnerships
Qualifications
  • 5+ years of software engineering, data engineering, or data platform experience
  • Strong proficiency in Python and modern software development practices, including git, testing, code review, CI/CD, and production deployment
  • Experience designing and operating production data pipelines, data warehouses, and data models at scale
  • Hands‑on experience with cloud platforms, preferably GCP, and technologies such as BigQuery, PostgreSQL, object storage, workflow orchestration, and containerized services
  • Strong understanding of data security, access control, data partitioning or siloing, audit logging, and managing sensitive or restricted datasets
  • Experience working with complex, heterogeneous datasets and building systems that make them reliable, discoverable, and usable
  • Ability to work independently, make sound technical decisions, and drive projects from ambiguous requirements to production systems
  • BS, MS, or PhD in Computer Science, Engineering, Data Science, Bioinformatics, or a related technical field, or equivalent practical experience
Preferences (but not required)
  • Experience with scientific, biological, clinical, genomic, laboratory, or high-throughput experimental data
  • Experience managing external partner, customer, or restricted-access datasets
  • Familiarity with data governance, lineage, metadata systems, schema registries, or data catalogs
  • Experience with research data systems, LIMS, ELNs, Benchling, or adjacent scientific platforms
  • Background working with ML, data science, computational biology, or cross-disciplinary technical teams
  • Interest in learning biology, gene editing, protein design, or machine learning concepts
What We Offer
  • High-growth opportunity with meaningful impact on the future of protein design
  • Competitive compensation package with equity participation
  • 401(k) with a strong employer match
  • Comprehensive benefits including health/dental/vision insurance
  • Generous PTO policy and commitment to work‑life balance
  • Professional development opportunities in a cutting‑edge field at the intersection of AI and biology

Profluent Bio, Inc is an equal opportunity employer promoting diversity and inclusion in the workspace. We do not discriminate on the basis of race, color, religion, marital status, age, national origin, ancestry, physical or mental disability, medical conditions, veteran status, sexual orientation, gender (including gender identity and gender expression), sex (which includes pregnancy, childbirth, and breastfeeding), genetic information, taking or requesting statutorily protected leave, or any other basis protected by law.

Work Authorization Requirement

Applicants must have ongoing work authorization in the United States that does not require employer sponsorship. Sponsorship will not be provided now or at any time in the future for this position.

Employment Eligibility Verification

Legal authorization to work in the United States is required. In compliance with federal law, all persons hired must verify their identity and work eligibility and complete the required employment verification form upon hire.

Hiring Salary Range

$170,000 — $220,000 USD

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Software Engineer, Data Platform
Senior Software Engineer, Data Platform

Profluent • Emeryville (CA)

On-site
USD 170,000 - 220,000
Competitive compensation package with equity participation
401(k) with a strong employer match
Generous PTO policy
+1
Machine Learning Research Engineer
Machine Learning Research Engineer

Profluent Bio • Emeryville (CA)

On-site
USD 200,000 - 330,000
Health insurance
Dental insurance
Vision insurance
+3
Scientist II, ML – Guided Protein Design Evaluation
Scientist II, ML – Guided Protein Design Evaluation

Profluent Bio • Emeryville (CA)

On-site
USD 147,000 - 180,000
Competitive compensation with equity participation
Comprehensive health benefits
Generous PTO policy
Biotech Project Manager, Early Discovery
Biotech Project Manager, Early Discovery

Profluent Bio • Emeryville (CA)

On-site
USD 115,000 - 150,000
401(k) with employer match
Comprehensive health benefits
Generous PTO
+1
Lead Bioinformatics Scientist, NGS
Lead Bioinformatics Scientist, NGS

Profluent Bio • Emeryville (CA)

On-site
USD 160,000 - 230,000
Competitive compensation package with equity
401(k) with employer match
Comprehensive health/dental/vision insurance
+2
Machine Learning Research Engineer
Machine Learning Research Engineer

Profluent • Emeryville (CA)

Hybrid
USD 200,000 - 330,000
Equity participation
401(k) match
Generous PTO
+1
Machine Learning Scientist, BioML
Machine Learning Scientist, BioML

jobr.pro • Emeryville (CA)

Hybrid
USD 200,000 - 330,000
Equity participation
401(k) with employer match
Health/dental/vision insurance
+2
Machine Learning Scientist, BioML
Machine Learning Scientist, BioML

Profluent Bio • Emeryville (CA)

On-site
USD 200,000 - 330,000
Competitive compensation package with equity participation
401(k) with employer match
Comprehensive health benefits
+1
Scientist I, Screening Tech Development
Scientist I, Screening Tech Development

Profluent • Emeryville (CA)

On-site
USD 127,000 - 142,000
Competitive compensation package with equity participation
Comprehensive health benefits
Generous PTO policy
Machine Learning Scientist, Reinforcement Learning
Machine Learning Scientist, Reinforcement Learning

Profluent • Emeryville (CA)

On-site
USD 200,000 - 330,000
Competitive compensation with equity participation
401(k) with strong employer match
Comprehensive health/dental/vision insurance
+2