Axle
Axle1d ago
New
USD 130000-150000/yr

Senior Data Scientist, AI Retrieval Systems

Remotesenior
Data ScientistData
3 views0 saves0 applied

Quick Summary

Key Responsibilities

Model biomedical knowledge for rare disease research. Ingest disease and phenotype ontologies and controlled vocabularies into PostgreSQL with a maintainable release and refresh path,

Requirements Summary

Bachelor’s degree in Data Science , Computer Science, Bioinformatics, Biomedical Informatics, or a related field . An advanced degree is preferred.

Technical Tools
Data ScientistData
(ID: 2026-3395)

 


What We Offer

~7 min read
100% Medical, Dental & Vision Coverage for Employees
Paid Time Off and Paid Holidays
401K match up to 5%
Educational Benefits for Career Growth
Employee Referral Bonus
Flexible Spending Accounts: Healthcare (FSA)
Parking Reimbursement Account (PRK)
Dependent Care Assistant Program (DCAP)
Transportation Reimbursement Account (TRN)
Model biomedical knowledge for rare disease research. Ingest disease and phenotype ontologies and controlled vocabularies into PostgreSQL with a maintainable release and refresh path, reconcile identifiers across sources, and work through term hierarchies to determine what is clinically relevant for a given condition.
Build retrieval-augmented services that ground everyday language in clinical concepts. Embed term labels, definitions, and synonyms, retrieve candidates, and have a model disambiguate against context before any value is committed.
Treat retrieval as a database problem. Tune keyword and vector search over large biomedical corpora, and be ready to defend the recall and latency trade-offs you choose.
Build the ranking and relevance layers that decide what surfaces first, including domain-aware weighting and graceful degradation when a condition falls outside curated coverage.
Deliver the interfaces where this work becomes visible to users, in Next.js, React, and TypeScript. This covers question and confirmation flows, result presentation, and live status for long-running pipelines.
Deploy continuously onto NIH on-premises and high-performance computing Kubernetes environments. Helm charts, StatefulSets, secrets, ingress, GPU scheduling for self-hosted inference, and scheduled jobs are all in scope, and you will partner with the operations teams that run those environments instead of standing up parallel cloud infrastructure.
Build the evaluation that tells us whether retrieval and concept mapping are good enough to rely on, and keep it running as a regression suite instead of a one-time measurement.
Log what the system does and why. Request identifiers, latency, errors, and which concept the system selected all need to be captured, so that staff can review an AI-assisted result instead of taking it on faith.
Work out what researchers, clinicians, and patient communities need, and turn it into data models, retrieval behavior, and interface design.
Write the work up. You will contribute to manuscripts, conference abstracts, and posters with NIH investigators, and you will be credited as an author on work you helped produce.
Bachelor’s degree in Data Science, Computer Science, Bioinformatics, Biomedical Informatics, or a related field. An advanced degree is preferred. We will consider equivalent professional experience in place of a degree.
At least 5 years building and operating production software or data systems. At least 2 of those years should involve shipping LLM-powered applications (agents, retrieval, or evaluation) that people depend on. We weigh depth in retrieval and applied LLM engineering more heavily than total years.
Experience building retrieval systems end to end, covering indexing, query construction, and measuring retrieval quality against real data.
Experience evaluating systems that have no single right answer, using golden sets, offline regression suites, or metrics such as Recall@K and MRR to decide whether a change was an improvement.
Experience with structured output and tool or function calling, meaning you have constrained a model to a typed schema and validated what came back.
Ability to own a service end to end, from schema design through deployment and operation.
Ability to obtain and maintain a Public Trust Security clearance.
Technical Skills:
Python, with FastAPI, Pydantic, and pytest.
PostgreSQL at depth, covering vector search (pgvector or equivalent), full-text search, embedding pipelines, indexing, and query tuning.
LLM application engineering: provider APIs and gateways, prompt and context design, structured generation, and tool use.
Data ingestion and transformation pipelines with a repeatable refresh path.
Containers and Kubernetes, enough to ship, debug, and operate a service on infrastructure you do not administer.
Working comfort in Next.js, React, and TypeScript.
Git-based collaboration and CI/CD in a shared codebase.
Biomedical ontologies and controlled vocabularies, including MONDO, HPO, UMLS, MeSH, and other OBO Foundry resources, along with comfort working through term hierarchies, synonyms, and cross references.
Grounding model output in a domain terminology through embedding-based retrieval plus model disambiguation, such as entity linking, concept normalization, or ontology alignment.
Helm, and deployment to on-premises or HPC Kubernetes environments.
Serving open-weight models in production with Ollama or vLLM behind a gateway such as LiteLLM, and work with domain embedding models such as MedCPT.
Background in rare disease, clinical genetics, or translational research.
Contributions to an open biomedical resource, standard, or consortium, such as OBO Foundry ontologies or GA4GH.
Published or presented work that explains your engineering to people who did not build it. Peer-reviewed papers, conference talks, preprints, technical blog posts, and public open source contributions all count.
Prior or current NIH experience.

Location & Eligibility

Where is the job
Worldwide
Fully remote, anywhere in the world
Who can apply
Same as job location

Listing Details

Posted
August 26, 2026
First seen
August 26, 2026
Last seen
August 27, 2026

Posting Health

Days active
0
Repost count
0
Trust Level
87%
Scored at
August 26, 2026

Signal breakdown

freshnesssource trustcontent trustemployer trust
Axle
Axle
greenhouse
Employees
5
Founded
2018
View company profile
Newsletter

Stay ahead of the market

Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.

A
B
C
D
Join 12,000+ marketers

No spam. Unsubscribe at any time.

AxleSenior Data Scientist, AI Retrieval SystemsUSD 130000-150000