afterquery
New
$210K – $450K • Offers Equity • Offers Bonus/yr

Research Scientist – Frontier Evaluations

United StatesUnited States·San Franciscofull-timemid
Data ScientistData
0 views0 saves0 applied

Quick Summary

Overview

About AfterQuery AfterQuery is an applied research lab curating data solutions for foundation model development.

Technical Tools
Data ScientistData

AfterQuery is an applied research lab curating data solutions for foundation model development. We serve every frontier AI lab with the mission of delivering the best data to power the best models. In doing so, we can make expertise that once took a lifetime to build available to anyone who needs it.

Our customers are the ones building the foundation models themselves and our work sits directly in the loop of how those systems improve. This is a rare opportunity to join a company at a defining moment in AI. We are YC's fastest unicorn, valued at $3.2 billion. We're based in San Francisco and backed by leading investors including Altos Ventures, BoxGroup, and Y Combinator and angels from Google DeepMind, OpenAI, Anthropic, Meta Superintelligence Labs, and Microsoft AI.

Massive Opportunity: We are YC's fastest unicorn valued at $3.2 billion and we're not slowing down.

Founding Impact: You will own and architect core infrastructure systems that power our platform from the ground up.

Equity & Growth: Competitive salary and meaningful equity. As we scale, you’ll have the opportunity to shape the engineering organization and lead major technical initiatives.

Strong Team: Our founding team has experience from Citadel Securities, Meta, Google, Silver Lake, and Morgan Stanley — work alongside world-class engineers and researchers.

AfterQuery is hiring Research Scientists to design and publish rigorous evaluations for frontier AI systems. The role spans agentic, coding, and safety evaluations, as well as expert-domain evaluations involving applied AI in healthcare, STEM, finance, and related fields. You will own evaluation development end to end and collaborate across disciplines to turn important capability gaps into rigorous public research.

Responsibilities

~1 min read
  • Lead the end-to-end design, validation, launch, and continuous improvement of frontier AI benchmarks.

  • Partner with researchers and domain experts to develop evaluations around meaningful model failures, gaps in existing coverage, and high-priority domains.

  • Analyze model capabilities and failure modes using rigorous experimental design and statistical methods.

  • Build reproducible evaluation systems, including harnesses, graders, and benchmark infrastructure.

  • Collaborate with researchers to post-train models and measure the resulting performance gains.

  • Communicate results through benchmark reports, technical articles, and research papers.

Requirements

~1 min read
  • Strong record of publishing benchmarks or research papers.

  • Clear technical communication and strong scientific writing skills.

  • Commitment to experimental rigor, including baselines, ablations, statistical validity, and contamination controls.

  • Ability to take an ambiguous evaluation question from initial scoping through a reproducible public release.

  • Depth in agentic, coding, and safety evaluations or applied machine learning in an expert domain.

  • PhD in a related technical field.

  • Research publications at leading conferences or peer-reviewed journals.

  • Interest in multidisciplinary research and the creativity to combine methods and insights from AI, engineering, science, and other expert domains.

What We Offer

~1 min read
Health Insurance: Medical, Vision, Dental
401(k) with Employer Match
Daily Meals: Daily UberEats Stipend
Monthly Wellness Stipend
Commute Covered

Location & Eligibility

Where is the job
San Francisco, United States
On-site at the office
Who can apply
US

Listing Details

Posted
September 10, 2026
First seen
September 11, 2026
Last seen
September 11, 2026

Posting Health

Days active
0
Repost count
0
Trust Level
63%
Scored at
September 11, 2026

Signal breakdown

freshnesssource trustcontent trustemployer trust
Newsletter

Stay ahead of the market

Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.

A
B
C
D
Join 12,000+ marketers

No spam. Unsubscribe at any time.

afterqueryResearch Scientist – Frontier Evaluations$210K – $450K • Offers Equity • Offers Bonus