Machine Learning Engineer
Quick Summary
About Arena Intelligence Arena Intelligence is the open platform for evaluating how AI models perform in the real world. Created by researchers from UC Berkeley’s SkyLab, our mission is to measure and advance the frontier of AI for real-world use.
Arena is the platform for evaluating how AI models perform in the real world. Founded by researchers from UC Berkeley's SkyLab, we're on a mission to measure and advance the frontier of AI for real-world use, and to build the foundation for everyone to understand, shape, and benefit from it.
Tens of millions of people use Arena each month to evaluate how frontier systems handle the work they actually do. The preferences they share power the most transparent, rigorous, and human-centered evaluations in AI. Leading AI labs, enterprises, and independent researchers rely on our work and open datasets to understand how models behave in real workflows: agentic coding, creative generation, professional productivity, and beyond. We go beyond leaderboards and decompose what human experience reveals about AI, so models advance toward the work people actually do.
We're a team of researchers, academics, builders, and creatives from UC Berkeley, Google, Stanford, and DeepMind. We seek truth, move fast, and value craftsmanship, curiosity, and impact over hierarchy. We're building a company where thoughtful, curious people from all backgrounds can do their best work together, in an office culture that radiates excellence, energy, and focus.
About the Role
~1 min readArena Intelligence is seeking a Senior Machine Learning Engineer to help scale and strengthen the core infrastructure that powers real-world AI evaluation. You’ll play a foundational role in shaping how we build, deploy, and improve our model benchmarking systems, working across data pipelines, inference APIs, and new evaluation methodologies. This is an opportunity to apply your technical expertise to a platform trusted by millions, and to help define how cutting-edge AI is assessed in the wild.
As one of the first ML engineers on the team, you’ll partner closely with researchers, engineers, and product leadership to turn new ideas into reliable systems. You’ll help us move fast while staying rigorous, improving reproducibility, scaling up to new modalities, and deepening our ability to understand and compare frontier models.
Architect and build what will become our core modeling for data and evaluation products
Own the full stack data, model training, and eval pipelines
Help grow a culture of feedback and rapid product iteration as we build new features as a tight-nit team
Conduct research into state-of-the-art evaluation methods and contribute to the long-term vision for a centralized, scalable evaluation platform.
Strong programming skills with the ability to work across the stack in a typical recommendation system or LLM stack
Experience in deep learning, language models or reward model training
Experience in working with LLM for fine tuning, prompt engineering, function calling etc
Self-motivated with a willingness to take ownership of tasks
A passion for shipping quality products
4+ years of industry experience or relevant projects
Solid understanding of statistics, and various tools and methodologies for evaluating uncertainty in a way that is specific to the given product being shipped
What We Offer
~1 min readLocation & Eligibility
Listing Details
- Posted
- December 18, 2025
- First seen
- May 6, 2026
- Last seen
- June 18, 2026
Posting Health
- Days active
- 43
- Repost count
- 0
- Trust Level
- 18%
- Scored at
- June 18, 2026
Signal breakdown
Please let arena know you found this job on Jobera.
4 other jobs at arena
View all →Explore open roles at arena.
Similar Machine Learning Engineer jobs
View all →Stay ahead of the market
Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.
No spam. Unsubscribe at any time.