causal
causal1d ago
New

Member of Technical Staff — Research Engineering, Evaluation

United StatesUnited States·San Franciscofull-timelead
OtherMember Of Technical Staff
0 views0 saves0 applied

Quick Summary

Key Responsibilities

comfortable building both backend pipelines and the frontend tools people read results in Owns deliverables end-to-end, from collecting

Technical Tools
OtherMember Of Technical Staff

Our mission is general causal intelligence; AI that is capable of (1) predicting the future and (2) identifying the actions to alter it.

 

To achieve this breakthrough, we are building a Large Physics foundation Model (LPM) because physical systems, unlike text or images, are governed by verifiable cause and effect. We believe that scaling on physics will enable an understanding of causality required to predict and control physical systems, starting with weather.

 

Our founding team has built and deployed AI against the physical world in robotics, drug discovery, and particle physics at institutions like DeepMind, Waymo, Cruise, Insitro, Nabla Bio, and CERN.

We look for research engineers who are excited to tackle unsolved problems. Progress is only as trustworthy as its measurement. As our model, data, and reasoning efforts multiply, every team needs to know, precisely and comparably, whether a change made the model better. Your mission is to build the central evaluation framework that the entire research organization runs on: the pipelines, the metrics, and the tools that turn results into shared understanding.

Responsibilities

~1 min read
  • Design and build a central, reusable evaluation framework that every model and every team runs through

  • Implement evaluation pipelines, benchmark suites, and baselines that make model quality measurable and comparable across efforts

  • Build the visualization and dashboard tools that turn raw results into shared, actionable understanding for the whole team

  • Establish sound statistical methodology for evaluation, so teams can distinguish real improvements from noise

  • Partner with research and domain teams to translate what "good" means in each domain into standardized, automated metrics

We value a relentless approach to problem-solving, rapid execution, and the ability to quickly learn in unfamiliar domains.

  • Strong software engineering skills and experience building data or evaluation pipelines at scale

  • Experience turning research or model outputs into metrics, benchmarks, and visualizations that teams rely on

  • Solid grasp of probability and statistics, with the judgment to design evaluations that measure what they claim to

  • Full-stack range: comfortable building both backend pipelines and the frontend tools people read results in

  • Owns deliverables end-to-end, from collecting requirements to autonomously driving execution

Location & Eligibility

Where is the job
San Francisco, United States
On-site at the office
Who can apply
US

Listing Details

Posted
July 20, 2026
First seen
July 20, 2026
Last seen
July 20, 2026

Posting Health

Days active
0
Repost count
0
Trust Level
52%
Scored at
July 20, 2026

Signal breakdown

freshnesssource trustcontent trustemployer trust
Newsletter

Stay ahead of the market

Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.

A
B
C
D
Join 12,000+ marketers

No spam. Unsubscribe at any time.

causalMember of Technical Staff — Research Engineering, Evaluation