Scientific Data Engineer

Lawrence Berkeley National Laboratory

San Francisco (CA)

Hybrid

USD 132,000 - 161,000

Full time

10 days ago
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Health benefits
Retirement plans
Winter holiday shutdown
Parental leave

Job summary

Lawrence Berkeley National Laboratory is seeking a Scientific Data Engineer in the Scientific Data Division. The role focuses on software and data engineering for multi-modal data modeling, with applications in omics, structural biology, and neurophysiology.

You will create tools for scientific data management, enable FAIR data practices, and support ML usage in collaborative, cross-disciplinary teams. The position emphasizes development of high-performance computing and cloud solutions,

Qualifications

  • Five years of related experience with a Bachelor's degree or three years with a Master's, or equivalent.
  • Experience developing software for data modeling or analysis in a scientific/research context.
  • Experience contributing to community-driven open source software.
  • Experience with data management, scientific data analysis, or machine learning.
  • Ability to work with domain scientists and translate requirements into designs.
  • Excellent oral and written communication skills.
  • Proven ability to work in a cross-disciplinary team.

Responsibilities

  • Design and develop user-friendly software packages for scientific data management and analysis.
  • Collaborate with domain experts to develop FAIR data models and management solutions.
  • Support machine learning use of biological data by making it well-structured and accessible.
  • Maintain open source software products including CI, testing, and releases.
  • Design and maintain high-performance computing and cloud solutions for visualization and analysis of complex data.
  • Develop ML/AI solutions for biological data in collaboration with scientists.
  • Train scientists and engineers in use of software at workshops and conferences.
  • Demonstrate good judgment in method selection and problem solving.
  • Network with senior internal and external personnel in related areas.

Skills

Team collaboration
Written communication
Oral communication
Cross-disciplinary teamwork
Problem solving

Education

Bachelor's in CS/Data Science
Master's degree
PhD in STEM

Tools

Git
GitHub
CI/CD
HPC
Cloud computing

Job description

Lawrence Berkeley National Laboratory is hiring a Scientific Data Engineer within the Scientific Data Division.

The Computational Biosciences Group has an immediate opening for a software and data engineer in the area of multi-modal data modeling and analysis with applications to bioscience research. You will develop new methods and software tools that enable scientific knowledge discovery using modern data management and machine learning technologies and advance the state-of-the-art in data-intensive analysis. Your projects will focus on the domains of omics/structural biology data and neurophysiology data. Under limited instruction, you will be part of an experienced team conducting R&D in the areas of FAIR data science, AI, and modern methods for data understanding. You will be working as part of a multi-disciplinary team composed of computer scientists, data scientists, and bioinformaticians. Please note this is a scientific software/data engineering position - it is not a pure machine learning or AI research position, and it is not a pure data science or analytics position.

You will:

Design and develop user-friendly software packages for scientific data management and analysis

Work with domain experts to develop FAIR data models (i.e., models of the structure organization of the data) and management solutions for bioscience applications

Support machine learning and AI use of biological data by making it well-structured, documented, and efficiently accessible

Maintain and manage open source software products, including managing development priorities, software releases, continuous integration, and testing

Design, implement and maintain high performance computing and cloud solutions for visualization and analysis of complex biological data

Develop machine learning and AI solutions for analysis of biological data in close collaboration with diverse teams of scientists

Train scientists and research software engineers in the use of the developed software products at workshops and conferences

Demonstrate good judgment in selecting methods and techniques for obtaining solutions.

Network with senior internal and external personnel in their own area of expertise.

We are looking for:

Typically requires a minimum of 5 years of related experience with a Bachelor's degree in computer science, data science, machine learning, bioinformatics, or equivalent; or 3 years and a Master's degree; or equivalent work experience designing and developing software for data modeling or analysis; or a PhD in a relevant STEM field

Demonstrated experience developing software in a scientific or research context, such as in a research group, a scientific user facility, or on a scientific software project

Demonstrated hands‑on experience in a production environment, developing scientific software, scientific data models, or scientific data pipelines

Experience contributing to community-driven open source software

Demonstrated experience in one or more of the following areas: data management, scientific data analysis, machine learning

Works well in a collaborative team environment

Demonstrated capability with the Git version control and continuous integration systems, such as GitHub or GitLab

Ability to work effectively with domain scientists whose expertise is outside computing, and to translate their requirements into technical designs.

Excellent oral and written communication skills.

Demonstrated ability to work effectively as part of a cross-disciplinary team.

Desired skills/knowledge:

Master's or PhD in Computer Science or related field, with 5 or more years of professional experience designing and developing scientific data modeling or analysis software

Experience working with modern scientific data formats and database systems, such as HDF5, Zarr, MongoDB, PostgreSQL, MySQL, and Redis

Experience with Neurodata Without Borders, LinkML, or similar software ecosystems

Experience working with large biological data, such as in the areas of neurophysiology, microbiology, genomics, or protein design

Experience designing or working with structured data models, schemas, ontologies, or data standards

Familiarity with FAIR data principles, persistent identifiers, provenance, and controlled vocabularies and ontologies

Experience preparing scientific datasets for use by machine learning pipelines or LLM-based agents

Experience working with cloud object storage, cloud computing, High-Performance Computing, data lakehouse architecture, or containerization.

Experience developing web-based graphical user interfaces (GUIs) or application programming interfaces (APIs) for scientific data analysis and management

We invest in our employees by offering a total rewards package you can count on:

Exceptional health and retirement benefits, including pension or 401K-style plans

A culture where you'll belong - we are invested in our teams!

In addition to accruing vacation and sick time, we also have a Winter Holiday Shutdown every year.

Parental bonding leave (for both mothers and fathers)

Additional information:

Appointment type: This is a full-time, 2 years, term appointment with the possibility of extension or conversion to Career appointment based upon satisfactory job performance, continuing availability of funds and ongoing operational needs.

Salary range: The expected salary for this position is $131,760 - $161,064, which fits into the full salary of $117,132 - $197,676 depending upon the candidate’s skills, knowledge, and abilities. This includes education, certifications, and years of experience.

Background check: This position is subject to a background check. Any convictions will be evaluated to determine if they directly relate to the responsibilities and requirements of the position. Having a conviction history will not automatically disqualify an applicant from being considered for employment.

Work modality: Work may be performed on-site, or hybrid. The primary location for this role is Lawrence Berkeley National Lab, 1 Cyclotron Road, Berkeley, CA. Work must be performed within the United States. A REAL ID or other acceptable form of identification is required to access Berkeley Lab sites (for more information)

Equal Employment Opportunity Employer: The foundation of Berkeley Lab is our Stewardship Values: Team Science, Service, Trust, Innovation, and Respect; and we strive to build community with these shared values and commitments. Berkeley Lab is an Equal Opportunity Employer. We heartily welcome applications from all who could contribute to the Lab's mission of leading scientific discovery, excellence, and professionalism. In support of our rich global community, all qualified applicants will be considered for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, age, protected veteran status, or other protected categories under State and Federal law.

Misconduct Disclosure Requirement: As a condition of employment, the final candidate who accepts an offer of employment will be required to disclose if they have been subject to any final administrative or judicial decisions within the last seven years determining that they committed any misconduct; or have filed an appeal of a finding of substantiated misconduct with a previous employer. For additional information

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Scientific Data Engineer
Scientific Data Engineer

Berkeley Lab • Berkeley (CA)

On-site
USD 132,000 - 161,000
Health and retirement benefits
Winter Holiday Shutdown
Parental bonding leave
+1
Scientific Data Engineer
Scientific Data Engineer

LBL • Town of Montana (WI)

Hybrid
USD 132,000 - 161,000
Scientific Data Engineer
Scientific Data Engineer

LBL • California (MO)

Hybrid
USD 132,000 - 161,000
Total rewards package
On-site or hybrid work environment
Collaborative interdisciplinary team
AI Predictive Biology Postdoctoral Fellow (KBase Project)
AI Predictive Biology Postdoctoral Fellow (KBase Project)

Lawrence Berkeley National Laboratory • Berkeley (CA)

On-site
USD 92,000 - 103,000
Exceptional health benefits
Generous paid time off
Inclusive culture
AI Predictive Biology Postdoctoral Fellow (KBase Project)
AI Predictive Biology Postdoctoral Fellow (KBase Project)

Berkeley Lab • Berkeley (CA)

On-site
USD 990 - 111,000
Generous paid time off
Culture of belonging
AI Software Developer (KBase Project)
AI Software Developer (KBase Project)

Berkeley Lab • Berkeley (CA)

On-site
USD 117,000 - 147,000
Exceptional health and retirement benefits
Winter Holiday shutdown
Parental bonding leave
Scientific Data Engineer - FAIR Bioscience Data & Tools
Scientific Data Engineer - FAIR Bioscience Data & Tools

LBL • Town of Montana (WI)

Hybrid
USD 132,000 - 161,000
AI-Readiness & Data Automation Postdoctoral Scholar
AI-Readiness & Data Automation Postdoctoral Scholar

Berkeley Lab • Berkeley (CA)

On-site
USD 73,000 - 100,000
Exceptional health benefits
Retirement benefits
Winter Holiday Shutdown
+1
Scientific Data Engineer: FAIR BioData & Tools
Scientific Data Engineer: FAIR BioData & Tools

Lawrence Berkeley National Laboratory • San Francisco (CA)

Hybrid
USD 132,000 - 161,000
Health benefits
Retirement plans
Winter holiday shutdown
+1
Postdoctoral Researcher – Scientific Machine Learning & Computational Chemistry
Postdoctoral Researcher – Scientific Machine Learning & Computational Chemistry

Lawrence Berkeley National Laboratory • San Francisco (CA)

Hybrid
USD 106,000 - 123,000