ForgeApply
Try it free

ForgeApply · Job listing

Senior Data Scientist, Agentic AI Systems

Axle

Remote · US

See all 42 open roles at Axle

Tailor your resume for this Axle job in about a minute.

ForgeApply rewrites your resume for this exact posting, then autofills the application on Axle's site with it. You review everything before it's sent. Free trial, no card required.

About this role

(ID: 2026-3432)

Axle is a bioscience and information technology company that offers advancements in translational research, biomedical informatics, and data science applications to research centers and healthcare organizations nationally and abroad. With experts in biomedical science, software engineering, and program management, we focus on developing and applying research tools and techniques to empower decision-making and accelerate research discoveries. We work with some of the top research organizations and facilities in the country including multiple institutes at the National Institutes of Health (NIH).

Benefits We Offer:

• 100% Medical, Dental & Vision Coverage for Employees

• Paid Time Off and Paid Holidays

• 401K match up to 5%

• Educational Benefits for Career Growth

• Employee Referral Bonus

• Flexible Spending Accounts:

• Healthcare (FSA)

• Parking Reimbursement Account (PRK)

• Dependent Care Assistant Program (DCAP)

• Transportation Reimbursement Account (TRN)

Axle is seeking a Senior Data Scientist, Agentic AI Systems to join our vibrant team supporting rare disease research at the National Institutes of Health (NIH). This is a Remote position within the United States.

Position Summary

Roughly 25 to 30 million people in the United States live with a rare disease. There are somewhere between 7,000 and 10,000 distinct rare conditions, and the large majority have no FDA-approved treatment.

Research on these conditions keeps running into the same obstacles. Published evidence for any one disease is thin and scattered across sources. The same clinical finding gets written down a dozen different ways depending on who recorded it. And the people with the most at stake, patients and their families, are usually the least equipped to read the specialist literature written about their own condition.

Large language models are well suited to this class of problem, and the research programs we support are investing in applying them carefully. In this role you will build the conversational AI systems that sit between a person and the research infrastructure. These are multi-turn workflows that ask sensible follow-up questions in plain language, capture the answers as validated structured data, and hand that structure off to the searches and analyses doing the scientific work. The emphasis is on systems people can rely on, which in practice means confirming every interpretation before it is saved and logging every automated decision so that it can be reviewed later.

This is a senior individual contributor position. You will own major components from design through deployment, work directly with NIH program staff, clinical geneticists, and rare disease information specialists, and help set the engineering standards for how AI gets applied on this team.

Core Responsibilities

• Build agentic AI systems for rare disease research workflows. This includes the conversation logic, the rules that decide when enough information has been gathered, and the confirmation steps that catch a misreading before it reaches anything downstream.

• Model outputs in Pydantic and use structured output and tool calling, so that every field a model produces is typed, validated , and traceable back to its source.

• Write, version, and regression test the prompts behind clinical and scientific reasoning tasks. Prompts and output schemas are treated as code here, with tests to match.

• Build evaluation for tasks that have no single right answer. Golden sets, offline regression suites, and model-based graders all have a place, and the results should be good enough to decide what ships .

• Keep multi-step LLM workflows responsive under load. This covers async design, concurrency limits, streaming partial results to the client, and timeout and failure handling that holds up in production.

• Log what the system does and why. Request identifiers, latency, errors, and the reasoning behind each automated choice all need to be captured, so that staff can review an AI-assisted result instead of taking it on faith.

• Work out what researchers, clinicians, and patient communities need, and turn it into data models and system behavior.

• Write the work up . You will contribute to manuscripts, conference abstracts, and posters with NIH investigators, and you will be credited as an author on work you helped produce.

Required Qualifications

• Bachelor’s degree in Data Science , Computer Science, Bioinformatics, Biomedical Informatics, or a related field . An advanced degree is preferred. We will consider equivalent professional experience in place of a degree.

• At least 5 years building and operating production software or data systems. At least 2 of those years should involve shipping LLM-powered applications (agents, retrieval, or evaluation) that people depend on. We weigh depth in agentic workflow engineering more heavily than total years.

• Experience with structured output and tool or function calling, meaning you have constrained a model to a typed schema and validated what came back.

• Experience evaluating systems that have no single right answer, using golden sets, offline regression suites, or model-based graders to decide whether a change was an improvement.

• Ability to own a service end to end, from schema design through deployment and operation.

• Ability to obtain and maintain a Public Trust Security clearance.

Technical Skills

• Python, with FastAPI , Pydantic , and pytest .

• LLM application engineering: provider APIs and gateways, prompt and context design, structured generation, tool use, and tracing.

• PostgreSQL, including work with embeddings or vector search alongside relational data.

• Asynchronous and concurrent Python, plus streaming results to a client.

• Containers and Kubernetes, enough to ship, debug, and operate a service on infrastructure you do not administer.

• Git-based collaboration and CI/CD in a shared codebase.

Preferred Skills

• A typed agent frame

Tailor your resume for this Axle role before you apply.

Tailor my resume for this job

Similar jobs

More like this: Data Scientist Jobs · Remote Data Scientist Jobs · Browse all jobs

Free ATS checker · How to Autofill Greenhouse Job Applications (Without Sending Junk)