J
Jobgether22h ago
New
USD 260000-320000/yr

Software Engineering Director, Agentic Evaluations

United StatesUnited StatesRemoteFull-timeexecutive
OtherEngineering Director
0 views0 saves0 applied

Quick Summary

Overview

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Software Engineering Director,

Technical Tools
OtherEngineering Director

This role will lead the engineering strategy and execution behind credible, scalable evaluations of AI agents offered by software vendors. You will own the core evaluation platform, from APIs and administration tools through workflow orchestration and production delivery. The position combines deep backend engineering with hands-on experience evaluating customer-facing AI agents using modern LLMs and agentic techniques. You will design reusable evaluation infrastructure that can scale across software categories and accelerate the launch of new benchmarks. Alongside building the platform, you will lead and mentor a team of engineers and serve as a technical authority on agent evaluation. This is an opportunity to shape how agent performance, accuracy, reliability, and policy adherence are measured in an AI-driven software ecosystem.

  • Lead the development and continuous improvement of evaluations that run against live software agents, identifying practical approaches for producing useful, credible, and repeatable results.
  • Own the full technology stack supporting the evaluation platform, including administrative tooling, APIs, evaluation workflows, orchestration, and production systems.
  • Design reusable evaluation primitives and architecture that can be applied across multiple software categories and reduce the effort required to onboard new agent types.
  • Turn recurring integration, onboarding, and maintenance processes into reusable AI skills, agents, or automated workflows that increase evaluation velocity.
  • Monitor emerging frameworks, methodologies, and industry practices for agent evaluation, identifying opportunities to improve measurement quality and distribution of results.
  • Partner closely with data science teams to develop proprietary benchmarks and richer approaches to evaluating agent performance.
  • Lead, mentor, and develop a core engineering team, providing technical direction and subject-matter expertise in AI agent evaluations.
  • Promote the effective use of evaluations across agent-focused products and initiatives throughout the wider organization.

Requirements

~2 min read
  • 10+ years of professional software development experience, primarily in backend or full-stack environments.
  • 2+ years of direct engineering management experience, including team leadership, mentoring, and technical direction.
  • Expert-level backend development skills in languages such as Python, Java/Kotlin, TypeScript/JavaScript, or Go, with strong experience in frameworks such as FastAPI or Node.js.
  • Hands-on experience building evaluations for customer-facing AI agents, including the use of agent trajectory traces and evaluation rubrics to measure task completion, accuracy, correctness, and/or policy adherence.
  • Direct experience using frontier models from providers such as OpenAI, Anthropic, or Google in LLM-as-a-judge applications.
  • Regular experience using coding-agent tools such as Claude Code, Codex, OpenCode, or Pi as part of a modern software development workflow.
  • Bachelor’s degree in computer science, engineering, or a related field.
  • Experience integrating agent tool use directly or through MCP servers is a plus.
  • Familiarity with benchmark frameworks such as STATE-Bench, tau2-bench, or similar evaluation approaches is desirable.
  • Experience using browser automation and agent tooling such as Playwright, browser-use, or Chrome DevTools MCP is advantageous.
  • Experience with durable execution and workflow platforms such as Temporal, DBOS, Cloudflare Workflows, or Vercel Workflows is helpful.
  • Familiarity with agent sandboxing technologies such as AWS E2B, Daytona, Cloudflare Containers, or Vercel Containers is also valuable.
  • Strong communication, technical leadership, mentoring, and cross-functional collaboration skills are essential.

What We Offer

~1 min read
✓Total earnings of approximately $260,000–$320,000, combining base salary and bonus.
✓Equity participation.
✓Performance-based bonus opportunities.
✓Fully remote position within the United States.
✓Flexible working environment designed to support distributed teams.
✓Unlimited paid time off.
✓Generous parental leave.
✓Comprehensive benefits designed to support employee well-being and flexibility.
✓Inclusive and diverse workplace with employee-led community initiatives and professional growth opportunities.

Location & Eligibility

Where is the job
United States
Remote within one country
Who can apply
US

Listing Details

Posted
September 26, 2026
First seen
September 27, 2026
Last seen
September 27, 2026

Posting Health

Days active
0
Repost count
0
Trust Level
68%
Scored at
September 27, 2026

Signal breakdown

freshnesssource trustcontent trustemployer trust
Newsletter

Stay ahead of the market

Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.

A
B
C
D
Join 12,000+ marketers

No spam. Unsubscribe at any time.

J
Software Engineering Director, Agentic EvaluationsUSD 260000-320000