Sr. Scientific Data Engineer, R&D Data Platform

Abbott Laboratories

United States

On-site

USD 140,000 - 190,000

Full time

35 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Abbott Laboratories' Science Office within Cancer Diagnostics seeks a Senior Scientific Data Engineer to lead end-to-end data solutions for cancer research and diagnostics. You will build reusable tools, pipelines, APIs, notebooks, and lightweight apps to organize, validate, transform, analyze, and share complex data.

You will partner with scientists and engineers, design scalable AWS-based platforms, mentor teammates, and drive data-quality standards across programs.

Qualifications

  • Bachelor's degree in a quantitative field.
  • Five or more years of relevant professional or three+ years with an advanced degree.
  • Advanced programming skills in Python.
  • Strong SQL skills with structured and semi-structured data.
  • Track record of building reusable, maintainable software.
  • Experience with data pipelines, Python packages, APIs, notebooks, or internal tools.
  • Hands-on experience using AWS for data processing.
  • Depth in AWS services and architecture to evaluate options.
  • Experience with quantitative research such as statistics or ML.
  • Experience cleaning/integrating data from multiple sources at scale.
  • Fluency with Git, automated testing, documentation, code review, and CI.
  • Ability to investigate ambiguous problems and deliver working solutions.

Responsibilities

  • Lead design and delivery of reusable tools and services for data ingestion, validation, transformation, documentation, discovery, and sharing.
  • Own one or more platform capability areas end to end, including design, implementation, adoption, and maintenance.
  • Develop maintainable solutions using Python and SQL, including packages, pipelines, APIs, notebooks, and lightweight apps.
  • Create self-service workflows for researchers with varying programming experience.
  • Partner with scientific teams to map studies to a technical roadmap.
  • Establish standards for organizing and harmonizing data from disparate sources.
  • Design automated data-quality and validation frameworks.
  • Improve dataset documentation, traceability, and discoverability.
  • Evaluate AWS services and translate requirements into architecture with DevOps partners.
  • Develop solutions using AWS data capabilities (S3, Athena, Glue, EMR, Lambda, SageMaker).
  • Prototype solutions for programs and generalize into reusable platform capabilities.
  • Provide technical leadership across projects: design reviews and trade-offs.
  • Mentor engineers through code review and pairing, raising standards.
  • Support hands-on data analysis when needed to accelerate research.
  • Apply quantitative judgment to data and analytical needs.
  • Use Spark or PySpark for distributed processing when appropriate.
  • Adhere to software engineering practices: version control, testing, documentation, CI.
  • Communicate technical concepts clearly to technical and leadership audiences.
  • Operate independently in evolving environments with ambiguous requirements.

Skills

Python programming
SQL
Data analysis
Software design
Git & CI

Education

Bachelor’s degree in a quantitative field

Tools

Spark / PySpark
AWS (S3, Glue, Athena, EMR, Lambda, SageMaker)
APIs / notebooks

Job description

Position Overview:

The Science Office within Abbott Cancer Diagnostics is seeking a Senior Scientific Data Engineer to lead the design and delivery of practical data solutions for cancer research and diagnostic development. This position sits at the intersection of software engineering, scientific data, and applied analysis. You will own significant pieces of our research data platform end to end, building reusable tools and workflows that help researchers organize, validate, discover, transform, analyze, and share complex data. These solutions may include Python packages, data pipelines, APIs, notebooks, workflow utilities, and lightweight web applications. This is not a traditional enterprise data warehousing position. The work centers on heterogeneous research data generated across scientific programs, including genomic, clinical, imaging, laboratory, and experimental data. Successful candidates will combine deep technical skills with an understanding of how quantitative research is conducted, and will be comfortable setting technical direction when a problem is still loosely defined. You will work closely with scientists, data scientists, bioinformaticians, software engineers, and R&D DevOps partners, and will often represent the team in cross-functional technical discussions. The ideal candidate is curious, resourceful, and able to turn ambiguous scientific needs into durable, reusable capabilities, while helping other engineers do the same.

Essential Duties and Responsibilities:
  • Lead the design and delivery of reusable tools and services for ingesting, validating, transforming, documenting, discovering, and sharing scientific data.
  • Own one or more platform capability areas end to end, including design, implementation, adoption, operational support, and long-term maintainability.
  • Develop maintainable solutions using Python and SQL, including software packages, data pipelines, APIs, notebooks, workflow utilities, and lightweight internal applications.
  • Create approachable, self-service workflows that allow researchers with varying levels of programming experience to prepare and share data consistently.
  • Partner directly with scientific teams to understand their studies, analytical workflows, data sources, and recurring technical challenges, and translate those needs into a prioritized technical roadmap.
  • Establish standards and reusable patterns for organizing and harmonizing data from disparate sources, including consistent structures, terminology, variable definitions, and mappings, and drive their adoption across teams.
  • Design automated data-quality and validation frameworks that identify missing, inconsistent, malformed, or unexpected data before it is used in downstream research.
  • Improve the documentation, traceability, and discoverability of scientific datasets, including clear descriptions of data content, origin, ownership, processing history, and intended use.
  • Evaluate AWS services and features for scientific data and analytical workflows. Translate research requirements into technical recommendations and partner with R&D DevOps teams on architecture, deployment patterns, and operational ownership.
  • Develop solutions that use AWS data and analytics capabilities, particularly Amazon S3 and related services such as Athena, Glue, EMR, Lambda, and SageMaker.
  • Prototype solutions for individual research programs and lead the work of generalizing successful approaches into reusable platform capabilities.
  • Provide technical leadership on designs that span multiple projects or teams: lead design reviews, document trade-offs and decisions, and align approaches with other engineers and technical leads.
  • Mentor engineers through code review, pairing, design feedback, and documentation, and raise the overall engineering standard of the team.
  • Support hands-on preparation and analysis of scientific data when needed to understand a problem, validate a solution, or accelerate a research effort.
  • Apply quantitative and scientific judgment when evaluating data, analytical requirements, and potential technical solutions.
  • Use Spark or PySpark when distributed processing is appropriate for large or computationally intensive datasets.
  • Apply and reinforce sound software-engineering practices, including version control, testing, code review, documentation, dependency management, continuous integration, and reproducible development.
  • Communicate technical concepts, design decisions, trade-offs, limitations, and project status clearly to technical, scientific, and leadership audiences.
  • Operate independently within an evolving environment: scope ambiguous problems, sequence the work, make defensible decisions when requirements are incomplete, and keep stakeholders informed.
Minimum Qualifications:
  • Bachelor’s degree in computer science, data science, engineering, statistics, mathematics, bioinformatics, computational science, or another relevant quantitative discipline.
  • Five or more years of relevant professional or applied research experience, or three or more years with an advanced degree in a relevant field.
  • Advanced programming skills in Python.
  • Strong SQL skills and experience working with structured and semi-structured data.
  • Demonstrated track record of building reusable, maintainable software that others depend on, rather than one-time scripts or analyses.
  • Experience designing and delivering several of the following: data pipelines, Python packages, APIs, analytical workflows, notebooks, or internal software tools.
  • Substantial hands‑on experience using AWS for data processing, analytics, scientific computing, or software development.
  • Sufficient depth in AWS services and architecture to evaluate technical options, justify design recommendations, and define infrastructure requirements with DevOps or cloud‑engineering partners.
  • Experience conducting or supporting quantitative research, such as statistical analysis, machine learning, computational modeling, or another data‑intensive research activity.
  • Experience cleaning, integrating, standardizing, or validating data from multiple sources at meaningful scale.
  • Fluency with software‑development practices such as Git, automated testing, technical documentation, code review, and continuous integration.
  • Demonstrated ability to investigate ambiguous problems, define an approach, and deliver a working solution with little guidance.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Sr. Scientific Data Engineer, R&D Data Platform
Sr. Scientific Data Engineer, R&D Data Platform

Abbott Laboratories • Northern (KY)

Hybrid
USD 78,000 - 156,000
Lead Scientific Data Platform Engineer (R&D)
Lead Scientific Data Platform Engineer (R&D)

Abbott Laboratories • United States

On-site
USD 140,000 - 190,000
Sr. Data Engineer
Sr. Data Engineer

Abbott • Madison (WI)

On-site
USD 78,000 - 156,000
Senior Scientific Data Engineer, Remote Data Platform
Senior Scientific Data Engineer, Remote Data Platform

Abbott Laboratories • Northern (KY)

Hybrid
USD 78,000 - 156,000
Associate Director, Principal Data Engineer
Associate Director, Principal Data Engineer

Scorpion Therapeutics • San Francisco (CA)

On-site
USD 180,000 - 260,000
Software Engineer Senior
Software Engineer Senior

Nationwide Children's Hospital • Columbus (OH)

On-site
USD 120,000 - 160,000
Scientific Data Architect
Scientific Data Architect

Insilico Search Partners • Boston (MA)

On-site
USD 100,000 - 130,000
Principal Scientist, Data Science – R&D, Therapeutics Development & Supply
Principal Scientist, Data Science – R&D, Therapeutics Development & Supply

Jobtailor • Spring House (PA)

On-site
USD 110,000 - 170,000
Sr Data Scientist
Sr Data Scientist

BioPharma Consulting JAD Group • Juncos (PR)

On-site
USD 120,000 - 180,000
6- month contract with extension
Administrative Shift
Data Engineer
Data Engineer

BioAgilytix • North Carolina

On-site
USD 120,000 - 160,000
Medical Insurance (HDHP with HSA)
Dental Insurance
Vision Insurance
+4