ForgeApply · Job listing
Microbiologist IV (Genomic Data Engineer)
Shinvestmentsllc
See all 48 open roles at Shinvestmentsllc →
Tailor your resume for this Shinvestmentsllc job in about a minute.
ForgeApply rewrites your resume for this exact posting, then autofills the application on Shinvestmentsllc's site with it. You review everything before it's sent. Free trial, no card required.
About this role
Great Hill Solutions, LLC is part of the Seneca Nation Group (SNG) portfolio of companies . SNG is Seneca Holdings' federal government contracting business that meets mission-critical needs of federal civilian, defense, and intelligence community customers. Our portfolio comprises multiple subsidiaries that participate in the Small Business Administration 8(a) program. To learn more about SNG, visit the website and follow us on LinkedIn .
Our team of talented individuals is what makes us successful. To support our team, we provide a balanced mix of benefits and programs. Your total rewards package includes competitive pay, benefits, and perks, flexible work-life balance, professional development opportunities, and performance and recognition programs. We offer a comprehensive benefits package that includes medical, dental, vision, life, and disability, voluntary benefit programs (critical illness, hospital, and accident), health savings and flexible spending accounts, and retirement 401K plan. One of our fundamental principles is to offer competitive health and welfare benefits to our team members, providing coverage and care for you and your family. Full-time employees working at least 30 hours a week on a regular basis are eligible to participate in our benefits and paid leave programs. We pride ourselves on our collaborative work environment and culture, which embraces our mission of providing financial and non-financial benefits back to the members of the Seneca Nation.
Great Hill is seeking a Microbiologist IV (Genomic Data Engineer) in Atlanta, GA.
The Microbiologist IV (Genomic Data Engineer) will provide scientific support to achieve the mission of the Coronavirus and Other Respiratory Viruses Division (CORVD). The role supports pathogen genomics, public health surveillance, outbreak detection, and epidemiological investigations through advanced genomic data engineering, integration, and analytics. The role also collaborates with multidisciplinary scientific teams, maintains technical documentation, prepares reports and scientific communications, and contributes to continuous improvement of data engineering and data management practices in support of public health objectives.
Job Description
• Develop, maintain, and optimize distributed data pipelines using Hadoop ecosystem tools (Hadoop Distributed File System, Spark, Hive, Impala).
• Manage large-scale ETL workflows involving genomic, epidemiological, and laboratory datasets to support bioinformatic workflows.
• Implement and optimize data validation, transformation, harmonization, and standardization workflows to ensure consistent, high-quality outputs.
• Ingest, harmonize, and manage genomic datasets from external repositories (e.g., NCBI GenBank, Sequence Read Archive) and maintain pipelines for routine updates and submissions.
• Work with genomic sequence files and associated metadata and integrate them into epidemiological and laboratory surveillance systems.
• Ensure appropriate handling of sensitive public health data and compliance with data governance expectations.
• Maintain reproducible workflows and version-controlled pipelines (e.g., Git) and prepare associated technical documentation.
• Collaborate with bioinformaticians, laboratory scientists, and epidemiologists to translate scientific questions into scalable engineered data workflows.
• Support development of analytical methods for outbreak detection and situational awareness, including Spark/SQL-based analysis.
• Document advanced data lineage, governance processes, or other high-level data management structures beyond required quality controls.
• Prepare reports, summaries, or scientific communication materials, and contribute to publications when appropriate.
• Be proficient in common programming or scripting languages, such as Python, Rust, Scala, and/or Bash
• Be present on site and attend weekly team meetings and provide updates on data engineering activities, pipeline performance, and ongoing tasks.
QUALIFICATIONS
Education and Experience:
• Bachelor's degree in Bioinformatics, Data Science, Genomics, Computational Biology or a related field.
• Master's degree is preferred in a relevant technical or scientific discipline.
Required Skils/Qualifications:
• Proficiency with Hadoop ecosystem technologies, including: Hadoop Distributed File System (HDFS), Apache Spark, Apache Hive, Apache Impala,
• Strong experience in data engineering, ETL development, and large-scale data integration.
• Experience with genomic, laboratory, epidemiological, or public health datasets.
• Ability to develop and optimize data validation, transformation, harmonization, and standardization processes.
• Experience ingesting and managing datasets from external genomic repositories such as NCBI GenBank and Sequence Read Archive (SRA).
• Proficiency working with genomic sequence files and associated metadata.
• Experience with version control systems, particularly Git.
• Knowledge of data governance, data quality management, and secure handling of sensitive health-related information.
• Proficiency in one or more programming and scripting languages such as: Python, Scala, Rust, Bash.
• Strong analytical, problem-solving, and technical documentation skills.
• Ability to collaborate effectively with multidisciplinary teams including bioinformaticians, epidemiologists, and laboratory scientists.
• Ability to work on-site and participate in regular team meetings and project updates.
Desirable Skills/Qualifications:
• Master's degree or higher in Bioinformatics, Computational Biology, Computer Science, Data Science, Public Health Informatics, or a related discipline.
• Experience supporting pathogen genomics and infectious disease surveillance programs.
• Advanced experience with Spark-based analytics and large-scale distributed computing environments.
• Familiarity with bioinformatics workflows, genomic analysis pipelines, and sequence data managemen
Salary insight
This posting doesn't disclose pay. Across 705 Atlanta jobs with disclosed salaries on ForgeApply, the median is $133k.
See full Data Engineer salary data for Atlanta →
Based on live postings with disclosed pay on ForgeApply; refreshed daily. Not an estimate of this employer's offer.
Tailor your resume for this Shinvestmentsllc role before you apply.
Tailor my resume for this jobSimilar jobs
- Staff Scientist - Genomics — Axle · Hamilton, MT
- Scientist II, Genomic Core — Xairatherapeutics · South San Francisco, California, United States
- Scientist II (Microbiology) — Fresenius Kabi · Wilson, NC
- Genomic Specialist — Veracyte · Remote
- Scientist I - Microbiology — Texas A&M University System · Canyon, TX
- Genomics Data Operations Engineer — Genomics · Oxford
- Genomics Technician — Axle · Rockville, MD
- Microbiologist I — IFF · Madison, WI
More like this: Data Engineer Jobs · Data Engineer Jobs in Atlanta · Browse all jobs
Free ATS checker · How to Autofill Greenhouse Job Applications (Without Sending Junk)