2h ago
New
$7,000 – $10,000/yr

Software Engineer | AI Training Data & Evals Lab

United StatesUnited StatesRemoteFull-timemid
OtherSoftware Engineer Ai
0 views0 saves0 applied

Quick Summary

Requirements Summary

Strong software engineering fundamentals with professional experience in Node.js and TypeScript. Strong coding ability in Python and/or Go. Demonstrated experience building, deploying,

Technical Tools
OtherSoftware Engineer Ai

This is a broad engineering role at the intersection of platform development, AI evaluation, and experimentation infrastructure.
You will build and operate production systems that researchers and operators rely on to develop and assess frontier AI capabilities.
Your work will include evaluation harnesses, backend services, data pipelines, training environments, and tools that make experimentation more reliable and repeatable.
You will have meaningful ownership from the start, with an expectation of delivering production improvements and becoming a trusted owner of core systems.
The environment is lean, async-first, and highly collaborative, with an emphasis on clear communication, sound technical judgment, and dependable execution.
You will work on practical engineering challenges closely connected to AI research while helping transform expert work into high-quality training and evaluation data.
This opportunity is well suited to a strong builder who enjoys autonomy, distributed systems, and working on infrastructure that directly influences how AI systems are evaluated and improved.

  • Build and maintain evaluation harnesses that measure the performance of AI models and agents on real-world tasks.
  • Improve evaluation reliability, coverage, and signal quality through better rubrics, task design support, and scoring approaches.
  • Develop tools that enable researchers and operators to run experiments efficiently without repeatedly rebuilding the same workflows.
  • Build and maintain APIs and backend services supporting human-in-the-loop workflows, task routing, and quality-control processes.
  • Improve data pipelines that transform expert work into structured training and evaluation datasets.
  • Strengthen system observability, scalability, and operational reliability through effective logging, metrics, monitoring, and debugging capabilities.
  • Write clear, maintainable production code and actively participate in code reviews, architecture discussions, and technical design decisions.
  • Document technical decisions and system behavior clearly so that other engineers and collaborators can build upon and operate the systems effectively.
  • Take ownership of core systems from development through production operation, with an expectation of delivering meaningful improvements within the first 30–90 days.

Requirements

~1 min read
  • Strong software engineering fundamentals with professional experience in Node.js and TypeScript.
  • Strong coding ability in Python and/or Go.
  • Demonstrated experience building, deploying, and owning production systems, including APIs, backend services, and data pipelines.
  • Solid understanding of distributed systems, scalability, reliability, and engineering trade-offs.
  • Experience working with AWS or GCP and modern infrastructure technologies such as containers and Kubernetes.
  • Proven track record of shipping and maintaining production systems that other people depend on, rather than working exclusively on prototypes.
  • Strong written communication skills and the ability to collaborate effectively in an asynchronous, distributed environment.
  • Comfortable taking ownership of ambiguous technical problems, making sound engineering decisions, and following projects through to production.
  • Experience with evaluation frameworks, experimentation platforms, or machine-learning tooling is a plus.
  • Experience with data pipelines, workflow orchestration, or internal platforms for research and operations teams is a plus.
  • Experience working in early-stage environments or high-ownership B2B SaaS and platform teams is a plus.

What We Offer

~2 min read
✓Full-time, fully remote position with a LATAM focus and meaningful overlap with U.S. time zones.
✓Compensation of $7,000–$10,000 USD per month, based on experience.
✓Significant ownership and opportunities to grow into larger systems, deeper technical leadership, and projects central to the organization’s growth.
✓Lean, async-first working environment focused on clear writing, sound judgment, and strong follow-through.
✓Opportunity to work on research-adjacent engineering challenges at the frontier of AI while building practical production platforms.
✓Direct impact on the training data and evaluation systems used by leading AI labs.
✓Structured hiring process including a practical take-home assignment, team review, technical screen, real-world work trial, and final offer stage.

Location & Eligibility

Where is the job
United States
Remote within one country
Who can apply
US

Listing Details

Posted
September 29, 2026
First seen
September 29, 2026
Last seen
September 29, 2026

Posting Health

Days active
0
Repost count
0
Trust Level
68%
Scored at
September 29, 2026

Signal breakdown

freshnesssource trustcontent trustemployer trust
Newsletter

Stay ahead of the market

Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.

A
B
C
D
Join 12,000+ marketers

No spam. Unsubscribe at any time.

Software Engineer | AI Training Data & Evals Lab$7k–$10k