Axle1d ago
New
New
USD 130000-150000/yr
Senior Data Scientist, AI Retrieval Systems
Remotesenior
Data ScientistData
3 views0 saves0 applied
Quick Summary
Key Responsibilities
Model biomedical knowledge for rare disease research. Ingest disease and phenotype ontologies and controlled vocabularies into PostgreSQL with a maintainable release and refresh path,
Requirements Summary
Bachelor’s degree in Data Science , Computer Science, Bioinformatics, Biomedical Informatics, or a related field . An advanced degree is preferred.
Technical Tools
Data ScientistData
(ID: 2026-3395)
What We Offer
~7 min read✓100% Medical, Dental & Vision Coverage for Employees
✓Paid Time Off and Paid Holidays
✓401K match up to 5%
✓Educational Benefits for Career Growth
✓Employee Referral Bonus
✓Flexible Spending Accounts:
Healthcare (FSA)
✓Parking Reimbursement Account (PRK)
✓Dependent Care Assistant Program (DCAP)
✓Transportation Reimbursement Account (TRN)
✓Model biomedical knowledge for rare disease research. Ingest disease and phenotype ontologies and controlled vocabularies into PostgreSQL with a maintainable release and refresh path, reconcile identifiers across sources, and work through term hierarchies to determine what is clinically relevant for a given condition.
✓Build retrieval-augmented services that ground everyday language in clinical concepts. Embed term labels, definitions, and synonyms, retrieve candidates, and have a model disambiguate against context before any value is committed.
✓Treat retrieval as a database problem. Tune keyword and vector search over large biomedical corpora, and be ready to defend the recall and latency trade-offs you choose.
✓Build the ranking and relevance layers that decide what surfaces first, including domain-aware weighting and graceful degradation when a condition falls outside curated coverage.
✓Deliver the interfaces where this work becomes visible to users, in Next.js, React, and TypeScript. This covers question and confirmation flows, result presentation, and live status for long-running pipelines.
✓Deploy continuously onto NIH on-premises and high-performance computing Kubernetes environments. Helm charts, StatefulSets, secrets, ingress, GPU scheduling for self-hosted inference, and scheduled jobs are all in scope, and you will partner with the operations teams that run those environments instead of standing up parallel cloud infrastructure.
✓Build the evaluation that tells us whether retrieval and concept mapping are good enough to rely on, and keep it running as a regression suite instead of a one-time measurement.
✓Log what the system does and why. Request identifiers, latency, errors, and which concept the system selected all need to be captured, so that staff can review an AI-assisted result instead of taking it on faith.
✓Work out what researchers, clinicians, and patient communities need, and turn it into data models, retrieval behavior, and interface design.
✓Write the work up. You will contribute to manuscripts, conference abstracts, and posters with NIH investigators, and you will be credited as an author on work you helped produce.
✓Bachelor’s degree in Data Science, Computer Science, Bioinformatics, Biomedical Informatics, or a related field. An advanced degree is preferred. We will consider equivalent professional experience in place of a degree.
✓At least 5 years building and operating production software or data systems. At least 2 of those years should involve shipping LLM-powered applications (agents, retrieval, or evaluation) that people depend on. We weigh depth in retrieval and applied LLM engineering more heavily than total years.
✓Experience building retrieval systems end to end, covering indexing, query construction, and measuring retrieval quality against real data.
✓Experience evaluating systems that have no single right answer, using golden sets, offline regression suites, or metrics such as Recall@K and MRR to decide whether a change was an improvement.
✓Experience with structured output and tool or function calling, meaning you have constrained a model to a typed schema and validated what came back.
✓Ability to own a service end to end, from schema design through deployment and operation.
✓Ability to obtain and maintain a Public Trust Security clearance.
✓Technical Skills:
✓Python, with FastAPI, Pydantic, and pytest.
✓PostgreSQL at depth, covering vector search (pgvector or equivalent), full-text search, embedding pipelines, indexing, and query tuning.
✓LLM application engineering: provider APIs and gateways, prompt and context design, structured generation, and tool use.
✓Data ingestion and transformation pipelines with a repeatable refresh path.
✓Containers and Kubernetes, enough to ship, debug, and operate a service on infrastructure you do not administer.
✓Working comfort in Next.js, React, and TypeScript.
✓Git-based collaboration and CI/CD in a shared codebase.
✓Biomedical ontologies and controlled vocabularies, including MONDO, HPO, UMLS, MeSH, and other OBO Foundry resources, along with comfort working through term hierarchies, synonyms, and cross references.
✓Grounding model output in a domain terminology through embedding-based retrieval plus model disambiguation, such as entity linking, concept normalization, or ontology alignment.
✓Helm, and deployment to on-premises or HPC Kubernetes environments.
✓Serving open-weight models in production with Ollama or vLLM behind a gateway such as LiteLLM, and work with domain embedding models such as MedCPT.
✓Background in rare disease, clinical genetics, or translational research.
✓Contributions to an open biomedical resource, standard, or consortium, such as OBO Foundry ontologies or GA4GH.
✓Published or presented work that explains your engineering to people who did not build it. Peer-reviewed papers, conference talks, preprints, technical blog posts, and public open source contributions all count.
✓Prior or current NIH experience.
Location & Eligibility
Where is the job
Worldwide
Fully remote, anywhere in the world
Who can apply
Same as job location
Listing Details
- Posted
- August 26, 2026
- First seen
- August 26, 2026
- Last seen
- August 27, 2026
Posting Health
- Days active
- 0
- Repost count
- 0
- Trust Level
- 87%
- Scored at
- August 26, 2026
Signal breakdown
freshnesssource trustcontent trustemployer trust
Salary
USD 130000-150000
per year
External application · ~5 min on Axle's site
Please let Axle know you found this job on Jobera.
3 other jobs at Axle
View all →Explore open roles at Axle.
Similar Data Scientist jobs
View all →Newsletter
Stay ahead of the market
Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.
A
B
C
D
No spam. Unsubscribe at any time.
