Senior Software Engineer – LLM Evaluation
Quick Summary
3+ years of professional software-engineering experience. Strong full-stack development capabilities and experience building scalable, production-grade software.
This part-time consulting opportunity is designed for experienced software engineers interested in advancing the evaluation of large language models.
You will curate high-quality code, develop technical solutions, and evaluate AI-generated software against real-world engineering standards.
The role spans multiple programming languages and covers the complete software-development lifecycle, from architecture and prototyping through production and maintenance.
You will design verification mechanisms and contribute to benchmarks that make AI evaluation more rigorous, consistent, and reproducible.
Your engineering expertise will help research teams identify model strengths, weaknesses, and recurring coding errors.
You will collaborate remotely with technical and research professionals on projects at the intersection of software engineering and AI.
The flexible contractor structure requires a minimum of 10 hours per week, with the possibility of working up to 40 hours depending on project needs.
-
Curate high-quality code examples and technical datasets for model training, benchmarking, and evaluation.
-
Develop accurate solutions to software-engineering tasks and correct or improve implementations across multiple programming languages.
-
Work with technologies such as Python, JavaScript, ReactJS, C/C++, Java, Rust, and Go as relevant to project assignments.
-
Evaluate AI-generated code for technical correctness, maintainability, efficiency, scalability, reliability, and adherence to professional engineering standards.
-
Identify implementation weaknesses, recurring coding errors, and patterns that reveal limitations in AI-generated software.
-
Provide clear, structured rationales explaining technical evaluation decisions and assessment outcomes.
-
Build agents and automated mechanisms capable of assessing code quality and verifying software solutions.
-
Design reliable checks that support consistent and reproducible evaluation across repeated engineering tasks.
-
Evaluate AI capabilities across the full software-development lifecycle, including prototyping, architecture, API design, production implementation, experimentation, launch, monitoring, and maintenance.
-
Assess model-generated technical reasoning and decisions against practical software-engineering expectations.
-
Collaborate with research and cross-functional technical teams to define evaluation strategies and improve coding benchmarks.
-
Contribute to datasets used for training and benchmarking while maintaining rigorous standards for quality and technical accuracy.
-
Help improve coding-focused evaluation systems through iterative analysis of model performance.
-
Complete all work without using confidential, proprietary, unreleased, employer-restricted, client-restricted, or otherwise protected code, datasets, architecture materials, or technical information belonging to any third party.
Requirements
~1 min read-
3+ years of professional software-engineering experience.
-
Strong full-stack development capabilities and experience building scalable, production-grade software.
-
Strong understanding of software architecture, system design, API design, and production implementation.
-
Deep knowledge of software development, debugging, code review, and code-quality assessment.
-
Demonstrated ability to review, troubleshoot, and improve complex software implementations.
-
Proficiency in one or more relevant programming languages, including Python, JavaScript, Java, C++, Rust, or related technologies.
-
Experience with ReactJS, C, Go, or additional programming languages is valuable depending on project requirements.
-
Familiarity with software monitoring, operational maintenance, and production reliability.
-
Ability to reason across the complete software-engineering lifecycle and evaluate technical decisions from development through ongoing operation.
-
Strong analytical and problem-solving skills, combined with a rigorous and detail-oriented approach to technical evaluation.
-
Excellent written and verbal communication skills, including the ability to produce concise and well-structured evaluation rationales.
-
Ability to distinguish between technically correct implementations and solutions that may introduce scalability, reliability, maintainability, or architectural concerns.
-
Comfortable working independently and collaborating remotely with research and technical teams.
-
Must be based in the United States, Canada, or an eligible Western European country.
-
Willingness to complete a required AI video interview as part of the application process.
What We Offer
~2 min readLocation & Eligibility
Listing Details
- Posted
- October 6, 2026
- First seen
- October 6, 2026
- Last seen
- October 7, 2026
Posting Health
- Days active
- 0
- Repost count
- 0
- Trust Level
- 68%
- Scored at
- October 7, 2026
Signal breakdown
Browse Similar Jobs
Stay ahead of the market
Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.
No spam. Unsubscribe at any time.