Computational Chemist (Machine Learning) I / II

Aralez Bio

Berkeley (CA)

On-site

USD 165,000 - 180,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Medical/dental/vision insurance
401k
Flexible Spending Account
Paid time off
Weekly catered lunch
Professional development opportunities
Annual company retreat

Job summary

Aralez Bio is seeking a Data Scientist to power the next phase of our R&D platform. You will design, deploy, and own ML models that predict chemical properties and guide de novo design of small molecules and peptides.

You’ll build scalable data pipelines and collaborate closely with laboratory scientists to translate experiments into actionable predictions. You will lead ML infrastructure development, leverage chemoinformatics tools, and clearly communicate results to technical and non‑technical

Qualifications

  • Ph.D. or Master’s in Computational Chemistry, Chemoinformatics, or a related field with 2–4 years of post‑grad experience.
  • Experience building and deploying ML/DL models to predict chemical properties for de novo design.
  • Proficiency with chemoinformatics toolkits (RDKit) and molecular descriptors, fingerprints, and data curation.
  • Fluent in Python and data science stacks; hands-on SQL and scalable data pipelines.
  • Strong communication and ability to work with laboratory scientists to translate questions into ML solutions.
  • Ability to lead initiatives independently and deliver structured data pipelines.
  • Experience with data visualization to communicate results to technical and non‑technical stakeholders.

Responsibilities

  • Design, build, and deploy machine learning models to predict chemical properties and support de novo design.
  • Own scalable data pipelines that ingest, curate, and prepare experimental screening data for ML.
  • Partner with laboratory scientists to translate questions into ML approaches and prioritize compounds for validation.
  • Develop and improve ML infrastructure, including training, evaluation, retraining, and monitoring.
  • Apply chemoinformatics tools and representations to engineer features and improve predictive accuracy.
  • Communicate insights through clear visualizations and presentations for diverse stakeholders.
  • Contribute to the ML platform’s technical direction with cross‑functional collaboration.

Skills

Python
Chemoinformatics
Data visualization
Communication
Independent work
ML model deployment
Small molecule design

Education

PhD/MS in Computational Chemistry

Tools

RDKit
REINVENT
SQL

Job description

About the Company

We envision a world where we continuously, sustainably, and affordably improve human life with synthetic biology. Our proprietary set of enzymes and our manufacturing approach enables the production of thousands of amino acids that were previously too difficult and expensive to make. These “noncanonical” amino acids that we make are already catalyzing the creation of revolutionary and innovative products that are good for people and the planet.

Our technology outperforms traditional manufacturing approaches by 10–100x across the board. We already sell products to Big 10 Pharmaceutical companies, and our product family includes the key components of multi-billion dollar blockbuster drugs such as Ozempic and Mounjaro.

We are a diverse, passionate, interdisciplinary team that finds joy in building something new that leaves the world better than we found it. We strive to always learn and improve and have a deep desire to capitalize on our creativity and be exceptional at what we do.

About the role

We're looking for a Data Scientist to help build the machine learning capabilities that will power the next phase of our R&D platform. In this role, you'll report to the Director of R&D and work side-by-side with laboratory scientists, applying machine learning and chemoinformatics to turn experimental data into testable predictions that accelerate scientific discovery. You'll have the opportunity to own models from initial concept through deployment, influence the technical direction of our ML infrastructure, and build scalable data pipelines that become foundational to the team's work.

If you're excited by applying cutting-edge machine learning to real-world scientific problems in a highly collaborative, interdisciplinary environment, this is a chance to have a meaningful impact on both our research strategy and the company's future growth.

This is a full-time, on-site role based out of Berkeley, CA.

What You’ll Do
  • Design, build, and deploy machine learning models to predict chemical properties and support de novo design of small molecules and peptides, taking models from concept through validation and production.
  • Own the development and maintenance of scalable data pipelines that ingest, curate, and prepare experimental screening data for machine learning applications, ensuring reproducibility and reliability.
  • Partner closely with laboratory scientists to translate experimental questions into machine learning approaches, generate actionable predictions, and prioritize new compounds or target areas for validation.
  • Develop and improve the team's machine learning infrastructure, including model training, evaluation, retraining, monitoring, and continuous improvement as new experimental data becomes available.
  • Apply chemoinformatics tools and molecular representations to engineer features, evaluate model performance, and improve predictive accuracy across R&D workflows.
  • Communicate insights and recommendations through clear visualizations and presentations, helping both technical and non-technical stakeholders understand model performance and scientific findings.
  • Contribute to the technical direction of the machine learning platform, collaborating with scientists and other technical team members to identify opportunities for new models, datasets, and workflows that accelerate research.
What You’ll Bring
  • Ph.D. or Master’s in Computational Chemistry, Chemoinformatics, or a related field and 2-4 years of post-graduate experience, with a strong foundation in chemical structure representation and molecular property prediction.
  • Proven track record in building and deploying machine learning (ML) and deep learning models to predict chemical characteristics and drive de novo design for chemical compounds with superior properties. Experience working with small molecule or peptide datasets.
  • High proficiency with using chemoinformatics toolkits (e.g. RDKit) and molecular descriptors, fingerprints, and structural data curation. Experience with structure or ligand-based tools (e.g. REINVENT) and molecular dynamics simulations is a plus
  • Fluent in Python and data science stacks, such as numpy, pandas, scipy. Hands on experience with SQL and managing data pipelines to ensure models are reproducible and scalable.
  • Strong communication and interpersonal skills. Able to maintain highly productive working relationships with laboratory scientists to ingest screening data, refine models, and recommend testable predictions.
  • Able to work independently to lead efforts and turn conceptual goals into structured data pipelines.
  • Experience with data visualization tools to communicate model results and predictions to technical and non-technical stakeholders.
  • Highly collaborative & comfortable with interdisciplinary communication: Success in this role depends on working closely with laboratory scientists and other technical partners to translate scientific questions into machine learning solutions and communicate results that drive research decisions.
  • Self-directed and results-oriented problem-solver: You'll be expected to independently identify opportunities, overcome technical challenges, and move projects from concept to implementation while delivering meaningful outcomes in a fast-paced startup environment.
  • Able to own and navigate complex and ambiguous tasks: Many of the problems you'll tackle won't have established solutions, requiring you to bring structure to uncertainty, make sound technical decisions, and adapt as new data and discoveries emerge.
Compensation

Aralez Bio’s salary range for this position is $165,000 to $180,000 per year, based on your performance and experience. In addition, your total rewards package will include equity and benefits.

Benefits at Aralez Bio
  • Medical / dental / vision insurance coverage
  • 401k
  • Flexible Spending Account
  • Paid time off
  • Weekly catered lunch
  • Professional Development opportunities
  • Annual company retreat

Aralez Bio is an equal opportunity employer and does not discriminate based on race, color, religion, marital status, age, national origin, ancestry, physical or mental disability, medical conditions, veteran status, sexual orientation, gender, sex, and any other group protected under federal, state, or local laws.

We celebrate diversity and are committed to creating an inclusive environment for all employees. Please reach out if an accommodation is needed. In compliance with federal law, all persons hired will be required to verify identity and eligibility to work in the United States and to complete the required employment eligibility verification form upon hire.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Computational Chemist (Machine Learning) I / II
Computational Chemist (Machine Learning) I / II

Aralez-Bio • Berkeley (CA)

On-site
USD 165,000 - 180,000
Medical / dental / vision insurance
401k
Flexible Spending Account
+4
Computational Chemist (Machine Learning) I / II
Computational Chemist (Machine Learning) I / II

Aralez Bio • South Carolina

On-site
USD 165,000 - 180,000
Medical / dental / vision insurance
401k
Flexible Spending Account
+4
Machine Learning Engineer - Computational Drug Discovery
Machine Learning Engineer - Computational Drug Discovery

5AM Ventures • City of Watertown (NY)

On-site
USD 140,000 - 200,000
Competitive compensation
Equity
Health benefits
+3
Machine Learning Engineer - Computational Drug Discovery
Machine Learning Engineer - Computational Drug Discovery

5AM Ventures • Watertown (MA)

On-site
USD 140,000 - 190,000
Equity
Health benefits
Paid time off
+2
Machine Learning Scientist/Senior Machine Learning Scientist - Synthesis Planning and Optimizat[...]
Machine Learning Scientist/Senior Machine Learning Scientist - Synthesis Planning and Optimizat[...]

Genentech • San Francisco (CA)

On-site
USD 147,000 - 274,000
Discretionary annual bonus
Comprehensive benefits package
Senior Machine Learning Scientist, Protein ML, AI for Drug Discovery (AIDD)
Senior Machine Learning Scientist, Protein ML, AI for Drug Discovery (AIDD)

Genentech • San Francisco (CA)

On-site
USD 167,000 - 311,000
Machine Learning Scientist/Senior Machine Learning Scientist – Synthesis Planning and Optimizat[...]
Machine Learning Scientist/Senior Machine Learning Scientist – Synthesis Planning and Optimizat[...]

NLP PEOPLE • San Francisco (CA)

On-site
USD 147,000 - 274,000
Machine Learning Scientist/Senior Machine Learning Scientist - Agents for Applied Small Molecul[...]
Machine Learning Scientist/Senior Machine Learning Scientist - Agents for Applied Small Molecul[...]

Genentech • San Francisco (CA)

On-site
USD 147,000 - 274,000
Applied ML Scientist (Staff / Principal)
Applied ML Scientist (Staff / Principal)

Genesis Therapeutics • San Mateo (CA)

On-site
USD 120,000 - 160,000
Competitive compensation
Comprehensive health benefits
Unlimited PTO
+1
Applied ML Scientist (Staff / Principal)
Applied ML Scientist (Staff / Principal)

Genesis Molecular AI • San Mateo (CA)

On-site
USD 180,000 - 260,000
Competitive compensation package
Comprehensive health benefits
401(k) plan
+4