Member of Technical Staff, Coding Agents
Quick Summary
About Handshake Handshake's mission is to organize expert human knowledge to advance the AI economy. Handshake AI works directly with frontier labs on their most consequential data, evaluation,
Handshake's mission is to organize expert human knowledge to advance the AI economy. Handshake AI works directly with frontier labs on their most consequential data, evaluation, and post-training challenges, building the systems that turn expert human knowledge into the data and evaluations that make frontier models better.
You will work alongside engineers, researchers, operators, and builders from organizations including Scale AI, Meta, Google, Amazon, xAI, Notion, and Palantir—and help build the systems that make expert human knowledge useful for advancing AI.
We are hiring a Member of Technical Staff to help define how frontier coding agents are benchmarked and improved. This is a broad, high-ownership role for researchers who build.
You will partner with the world’s top AI researchers and thousands of software engineers to develop new benchmarks, reward and verifier systems, agent-evaluation methodologies, and data-quality techniques, centered around applications of AI for Coding. You will work on questions at the center of frontier AI progress: what should be measured, how to design evaluations that reflect real capability, how to create high-signal feedback, and how to build the environments and data systems that make those answers actionable.
Early members of the team will have unusual influence over our technical direction, operating culture, and the open-source software, benchmarks, and research products we build. We care more about demonstrated research capability, technical judgment, and a builder's mindset than a specific title, degree, or career path.
Location: San Francisco preferred; we are open to exceptional candidates in other locations.
Responsibilities
~1 min read- →
Create/publish Coding benchmarks that challenge state-of-the-art frontier agents.
- →
Understand/shape the frontier of coding agents and where the space is headed, publishing studies of interesting/impactful findings.
- →
Help the world’s top AI Labs improve their models for coding.
- →
Publicly contribute to the field through benchmarks, open-source tools, research, and technical writing.
You’re extremely passionate about building software with agents, and regularly tinkering with the latest coding agents.
Comfort operating in an ambiguous, fast-moving environment with substantial ownership.
At least one of the following:
Previous experience working on a Coding benchmark.
Published research on AI for Software / Code-Generation at NeurIPS, ICLR, ICML, COLM, etc.
Extensive software engineering experience or strong GitHub profile of OSS contributions, along with deep knowledge of agent harnesses.
What We Offer
~1 min readLocation & Eligibility
Listing Details
- Posted
- October 9, 2026
- First seen
- October 9, 2026
- Last seen
- October 9, 2026
Posting Health
- Days active
- 0
- Repost count
- 0
- Trust Level
- 66%
- Scored at
- October 9, 2026
Signal breakdown
4 other jobs at
View all →Browse Similar Jobs
Stay ahead of the market
Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.
No spam. Unsubscribe at any time.