Applied AI Engineer
Quick Summary
Judgment Labs builds infrastructure for Agent Behavior Monitoring (ABM). While traditional observability focuses on logging exceptions and latency, our ABM surfaces behavioral anomalies such as instruction drifts and context retrieval loss in scaled production environments.
We are looking for Research Engineers to build AI systems that use agent interaction data to help us understand how agents behave, evaluate them at scale, and improve them through learning and feedback. Your research will not live on a whiteboard.
You identify with at least one of the following: You care about data quality, evaluation, and benchmarking, and are comfortable working hands-on with messy data You have experience building agent systems and working with them in real-world or…
We are looking for Applied AI Engineers to build AI systems that use agent interaction data to understand how agents behave, evaluate them at scale, and improve them through learning and feedback.
Your research will not live on a whiteboard. You’ll work directly with real-world agent data, apply frontier methods in production, and see your work ship into the product. By making agent behavior measurable and debuggable, your systems will support teams deploying agents across finance, legal, operations, and other high-stakes workflows. You will own projects end-to-end, with significant autonomy, and work closely with the team to build self-improving agent systems.
Responsibilities
~1 min read- →
Build AI systems to aggregate, index, and analyze large-scale long-running agent interaction data in order to extract meaningful signals
- →
Design and implement post-training and optimization workflows to improve agents, both internally and for customers
- →
Build agent platform infrastructure, including orchestration, runtimes, and developer tools that help teams define, test, deploy, and iterate on complex agent workflows
- →
Build internal tools and infrastructure that support rapid experimentation, analysis, and training
- →
Work closely with product to integrate agents into customer-facing workflows
- →
Collaborate with external companies and research partners on frontier AI research
Every hire clears three bars, no exceptions:
Agency. You are intellectually curious, self-directed, and stay up to date with the latest research, blogs, trends, and ideas.
Depth of thought. You can reason clearly about abstract systems, and ideally have experience working on agents, RL, or the infrastructure that supports them.
Ownership. You own outcomes, not just tasks. You use freedom to experiment responsibly, make business-driven decisions, and focus first on work that moves the company forward.
More specifically, you should bring strength in at least one of the following areas:
Data quality, evaluation, benchmarking, and hands-on work with messy production data
Agent systems built or evaluated in real-world or production settings
Reinforcement learning, post-training, agents, or machine learning fundamentals
Infrastructure and systems work across training, data pipelines, evaluation, or model serving
Translating research into product while balancing customer constraints, technical tradeoffs, and business impact
Turning ambiguous problems into clear, well-designed plans
Location & Eligibility
Listing Details
- Posted
- January 11, 2026
- First seen
- May 7, 2026
- Last seen
- September 26, 2026
Posting Health
- Days active
- 142
- Repost count
- 0
- Trust Level
- 15%
- Scored at
- September 26, 2026
Signal breakdown
Similar Applied Ai Engineer jobs
View all →Stay ahead of the market
Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.
No spam. Unsubscribe at any time.