Why This Role Stands Out
This leadership role offers a unique opportunity to shape the future of AI by directly impacting the evaluation and improvement of cutting-edge LLMs, with significant potential for long-term engagement. If you are a seasoned software engineer with a passion for AI and rigorous evaluation, you will thrive in this position, contributing to groundbreaking work in a hybrid environment. Apply now to be at the forefront of AI innovation!
Quick Overview
Seniority
Leader
Work mode
Hybrid
Location
Madrid, MD, United States
Posted
14 hours ago
Job Description
This is a contracting engagement - initially 6 months - with potential for long term engagement.
Location: Paris or London-based preferred; alternatively Europe remote for strong candidates
We are building and evaluating state-of-the-art large language models (LLMs) and are looking for experienced software engineers to join our evaluation and annotation team. This role sits at the intersection of real-world software engineering, model evaluation, and applied AI , and is critical to improving model reliability, reasoning, and code quality.
You will design challenging coding tasks, evaluate model outputs against rigorous benchmarks, identify failure modes, and contribute to reinforcement learning and model improvement workflows.
This is not a junior annotation role. We are looking for practitioners with deep hands-on coding experience who can think like both an engineer and an evaluator.
What You'll Do
Location: Paris or London-based preferred; alternatively Europe remote for strong candidates
We are building and evaluating state-of-the-art large language models (LLMs) and are looking for experienced software engineers to join our evaluation and annotation team. This role sits at the intersection of real-world software engineering, model evaluation, and applied AI , and is critical to improving model reliability, reasoning, and code quality.
You will design challenging coding tasks, evaluate model outputs against rigorous benchmarks, identify failure modes, and contribute to reinforcement learning and model improvement workflows.
This is not a junior annotation role. We are looking for practitioners with deep hands-on coding experience who can think like both an engineer and an evaluator.
What You'll Do
- Evaluate coding tasks involving software vulnerabilities, exploit verification, and security patches.
- Create high-quality coding prompts and reference answers (benchmark-style, e.g. SWE-Bench-like problems).
- Evaluate LLM outputs for code generation, refactoring, debugging, and implementation tasks.
- Identify and document model failures, edge cases, and reasoning gaps.
- Perform head-to-head evaluations between private LLMs (Mistral-based) and leading external models.
- Build or configure coding environments to support evaluation and reinforcement learning (RL).
- Follow detailed annotation and evaluation guidelines with high consistency.
- 5+ years of professional software development experience .
- Strong Python skills (required).
- Knowledge of at least one additional programming language (bonus).
- Experience with professional code review, coding annotation, LLM/code evaluation, or benchmark design is a plus, but not required.
- Hands-on experience with vulnerability research, exploit reproduction or verification , or i mplementing, backporting, or validating security patches .
- Proven ability to apply structured evaluation criteria and write clear technical feedback.
- Fluent in English (written and spoken) .
- Team lead or mentoring experience is a strong plus.
- Work hands-on with cutting-edge LLMs.
- Apply real-world engineering judgment to model evaluation and improvement.
- High-impact, technical work with a focused, senior team.
Similar jobs
- DW
DevSecOps Engineer
NewDark Wolf Solutions
Tampa, Florida🇺🇸$160k - $180k/yrHybrid2 minutes agoSAFeAgileEngineering - AS
Controls Engineer
Apex Systems
North Reading, MA🇺🇸$70 - $120/hrOn-site1 week agoEmbedded SystemsRoboticsSCADAEngineering - BR
Principal Coding Annotator / LLM Evaluation Engineer
NewBraintrust
Barcelona, CT🇺🇸Hybrid14 hours agoEngineering - AT
Quality Engineer
Advantage Technical
Arden Hills, MN🇺🇸$37/hrHybrid6 days agoEngineering - ES
Integration / Interface Engineer (HL7 / FHIR / Mirth Connect Specialist)
NewEPAM Systems
United States🇺🇸Hybrid14 hours agoEngineering - B&
Senior Electrical Engineering Technician - Substation P&C Designer
NewBlack & Veatch
Overland Park, KS🇺🇸$96.9k - $161.8k/yrHybrid14 hours agoAutoCADBIMCAD+1Engineering