Microbiologist IV (Genomic Data Engineer)

Seneca Holdings

Atlanta (GA)

On-site

USD 120,000 - 170,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Great Hill Solutions, LLC seeks a Microbiologist IV (Genomic Data Engineer) in Atlanta, GA to support CORVD objectives. You will develop scalable genomic data pipelines, manage ETL for diverse datasets, and collaborate with epidemiologists and lab scientists on outbreak analytics.

The role emphasizes data governance, reproducible workflows, and documentation, with on-site work and regular team updates in a mission-focused environment.

Qualifications

  • Proficiency with Hadoop ecosystem technologies (HDFS, Spark, Hive, Impala).
  • Experience with genomic, laboratory, epidemiological, or public health datasets.
  • Developing and optimizing data validation, transformation, harmonization, and standardization processes.
  • Ingesting and managing datasets from external genomic repositories like GenBank and SRA.
  • Version control with Git and reproducible data pipelines.

Responsibilities

  • Develop, maintain, and optimize distributed data pipelines using Hadoop ecosystem tools.
  • Manage large-scale ETL workflows for genomic and epidemiological datasets.
  • Ingest and harmonize genomic datasets from external repositories and maintain update pipelines.
  • Collaborate with bioinformaticians, epidemiologists, and lab scientists to translate questions into scalable data workflows.
  • Prepare technical reports and documentation; ensure data governance and secure handling of health data.

Skills

Data engineering
ETL development
Large-scale data integration
Genomic datasets
Git
Data governance

Education

Bachelor's degree in Bioinformatics, Data Science, Genomics, Computational Biology or related field
Master's degree preferred

Tools

HDFS
Apache Spark
Apache Hive
Apache Impala
Git

Job description

Great Hill Solutions, LLC is part of the Seneca Nation Group (SNG) portfolio of companies . SNG is Seneca Holdings' federal government contracting business that meets mission-critical needs of federal civilian, defense, and intelligence community customers. Our portfolio comprises multiple subsidiaries that participate in the Small Business Administration 8(a) program. To learn more about SNG, visit the website and follow us on LinkedIn . Our team of talented individuals is what makes us successful. To support our team, we provide a balanced mix of benefits and programs. Your total rewards package includes competitive pay, benefits, and perks, flexible work-life balance, professional development opportunities, and performance and recognition programs. We offer a comprehensive benefits package that includes medical, dental, vision, life, and disability, voluntary benefit programs (critical illness, hospital, and accident), health savings and flexible spending accounts, and retirement 401K plan. One of our fundamental principles is to offer competitive health and welfare benefits to our team members, providing coverage and care for you and your family. Full-time employees working at least 30 hours a week on a regular basis are eligible to participate in our benefits and paid leave programs. We pride ourselves on our collaborative work environment and culture, which embraces our mission of providing financial and non-financial benefits back to the members of the Seneca Nation.

Great Hill is seeking a Microbiologist IV (Genomic Data Engineer) in Atlanta, GA. The Microbiologist IV (Genomic Data Engineer) will provide scientific support to achieve the mission of the Coronavirus and Other Respiratory Viruses Division (CORVD). The role supports pathogen genomics, public health surveillance, outbreak detection, and epidemiological investigations through advanced genomic data engineering, integration, and analytics. The role also collaborates with multidisciplinary scientific teams, maintains technical documentation, prepares reports and scientific communications, and contributes to continuous improvement of data engineering and data management practices in support of public health objectives.

Job Description
  • Develop, maintain, and optimize distributed data pipelines using Hadoop ecosystem tools (Hadoop Distributed File System, Spark, Hive, Impala).
  • Manage large-scale ETL workflows involving genomic, epidemiological, and laboratory datasets to support bioinformatic workflows.
  • Implement and optimize data validation, transformation, harmonization, and standardization workflows to ensure consistent, high-quality outputs.
  • Ingest, harmonize, and manage genomic datasets from external repositories (e.g., NCBI GenBank, Sequence Read Archive) and maintain pipelines for routine updates and submissions.
  • Work with genomic sequence files and associated metadata and integrate them into epidemiological and laboratory surveillance systems.
  • Ensure appropriate handling of sensitive public health data and compliance with data governance expectations.
  • Maintain reproducible workflows and version-controlled pipelines (e.g., Git) and prepare associated technical documentation.
  • Collaborate with bioinformaticians, laboratory scientists, and epidemiologists to translate scientific questions into scalable engineered data workflows.
  • Support development of analytical methods for outbreak detection and situational awareness, including Spark/SQL-based analysis.
  • Document advanced data lineage, governance processes, or other high-level data management structures beyond required quality controls.
  • Prepare reports, summaries, or scientific communication materials, and contribute to publications when appropriate.
  • Be proficient in common programming or scripting languages, such as Python, Rust, Scala, and/or Bash Be present on site and attend weekly team meetings and provide updates on data engineering activities, pipeline performance, and ongoing tasks.
QUALIFICATIONS Education and Experience:
  • Bachelor's degree in Bioinformatics, Data Science, Genomics, Computational Biology or a related field.
  • Master's degree is preferred in a relevant technical or scientific discipline.
Required Skils/Qualifications:
  • Proficiency with Hadoop ecosystem technologies, including: Hadoop Distributed File System (HDFS), Apache Spark, Apache Hive, Apache Impala, Strong experience in data engineering, ETL development, and large-scale data integration.
  • Experience with genomic, laboratory, epidemiological, or public health datasets.
  • Ability to develop and optimize data validation, transformation, harmonization, and standardization processes.
  • Experience ingesting and managing datasets from external genomic repositories such as NCBI GenBank and Sequence Read Archive (SRA).
  • Proficiency working with genomic sequence files and associated metadata.
  • Experience with version control systems, particularly Git.
  • Knowledge of data governance, data quality management, and secure handling of sensitive health-related information.
  • Proficiency in one or more programming and scripting languages such as: Python, Scala, Rust, Bash.
  • Strong analytical, problem-solving, and technical documentation skills.
  • Ability to collaborate effectively with multidisciplinary teams including bioinformaticians, epidemiologists, and laboratory scientists.
  • Ability to work on-site and participate in regular team meetings and project updates.
Desirable Skills/Qualifications:
  • Master's degree or higher in Bioinformatics, Computational Biology, Computer Science, Data Science, Public Health Informatics, or a related discipline.
  • Experience supporting pathogen genomics and infectious disease surveillance programs.
  • Advanced experience with Spark-based analytics and large-scale distributed computing environments.
  • Familiarity with bioinformatics workflows, genomic analysis pipelines, and sequence data management.
  • Experience with analytical methods related to outbreak detection and situational awareness.
  • Knowledge of public health surveillance systems and laboratory information management systems.
  • Experience creating and maintaining data lineage documentation and enterprise data governance frameworks.
  • Experience contributing to technical reports, scientific publications or peer-reviewed research.
  • Familiarity with cloud-based data platforms and modern data engineering practices.
  • Strong communication skills with the ability to translate scientific and public health requirements into scalable technical solutions.

Equal Opportunity Statement: Seneca Holdings provides equal employment opportunities to all employees and applicants without regard to race, color, religion, sex/gender, sexual orientation, national origin, age, disability, marital status, genetic information and/or predisposing genetic characteristics, victim of domestic violence status, veteran status, or other protected class status. This policy applies to all terms and conditions of employment, including, but not limited to, hiring, placement, promotion, termination, layoff, recall, transfer, leave of absence, compensation and training. The Company also prohibits retaliation against any employee who exercises his or her rights under applicable anti-discrimination laws. Notwithstanding the foregoing, the Company does give hiring preference to Seneca or Native individuals. Veterans with expertise in these areas are highly encouraged to apply.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

RCMI Senior Solution Technical Expert (SSTE), Data Analyst
RCMI Senior Solution Technical Expert (SSTE), Data Analyst

Seneca Holdings • United States

On-site
USD 80,000 - 100,000
Comprehensive benefits package
Flexible work-life balance
Professional development opportunities
Genomic Data Engineer for Public Health Analytics
Genomic Data Engineer for Public Health Analytics

Seneca Holdings • Atlanta (GA)

On-site
USD 120,000 - 170,000
Senior Data Scientist
Senior Data Scientist

Seneca Holdings • Buffalo (NY)

On-site
USD 120,000 - 180,000
Senior Data Scientist
Senior Data Scientist

Seneca Holdings • United States

On-site
USD 120,000 - 160,000
Medical insurance
Dental insurance
Vision insurance
+2
Bioinformatician
Bioinformatician

Precise Software Solutions, Inc. • Atlanta (GA)

Hybrid
USD 70,000 - 90,000
Comprehensive Health Benefits
Flexible Spending Accounts
Retirement Plan with matching
+2
Junior Data & Insights Specialist
Junior Data & Insights Specialist

Seneca Holdings • United States

On-site
USD 55,000 - 75,000
Medical benefits
Dental benefits
Vision benefits
+3
Bioinformatician 1
Bioinformatician 1

Millipore Corporation • Rockville (MD)

On-site
USD 88,000 - 133,000
Health insurance
PTO
Retirement contributions
+1
Junior Data & Insights Specialist
Junior Data & Insights Specialist

Seneca Holdings • Chantilly (VA)

On-site
USD 60,000 - 75,000
Comprehensive benefits package
Flexible work-life balance
Professional development opportunities
Computational Biologist Genomics
Computational Biologist Genomics

Cherokee Federal • Frederick (MD)

On-site
USD 75,000
Journeyman Data & Insights Specialist
Journeyman Data & Insights Specialist

Seneca Holdings • Chantilly (VA)

Hybrid
USD 80,000 - 100,000
Health insurance
Dental insurance
401K retirement plan
+1