AI Benchmark Engineer | Native Language Specialist - French (France) - Remote

France (Remote)CompetitiveOnsite0 applicants

About this role

About The Opportunity

We are building a rigorous, verifiable evaluation suite of Terminal-Bench tasks designed to test the limits of large language models on multilingual software challenges. Our goal is to measure multilingual robustness across prompt language effects, non-English data processing, and complex locale/encoding edge cases in terminal workflows.

We are seeking experienced native-speaking software engineers to design, build, and validate these benchmarks. You will create high-signal, high-quality tasks that genuinely test a model's ability to handle multilingual environments without relying on English translation crutches.

Note this is a remote, freelance opportunity

What You’ll Deliver

Task Engineering: Evaluating Coding Agents.

Asset Creation: Build realistic task environments using datasets and files in your native language. Crucially, these assets must remain in the target language to genuinely measure multilingual handling.

Prompting & Translation: finding failure points where AI does not work, in your native language

Implementation & Verification: Support the development of robust solutions (reference implementations) and write highly reliable, deterministic verifier scripts (using rubric-based judging only when strictly necessary).

Calibration & Execution: Analyze execution logs and calibrate task difficulty (Easy to Very Hard) using standard Terminal-Bench run configurations against various model tiers (Haiku, Sonnet, Opus).

Quality Assurance: Participate in a rigorous, 4-layer

Responsibilities

  • Task Engineering: Evaluating Coding Agents.
  • Asset Creation: Build realistic task environments using datasets and files in your native language. Crucially, these assets must remain in the target language to genuinely measure multilingual handling.
  • Prompting & Translation: finding failure points where AI does not work, in your native language

Requirements

  • Implementation & Verification: Support the development of robust solutions (reference implementations) and write highly reliable, deterministic verifier scripts (using rubric-based judging only when strictly necessary).
  • Calibration & Execution: Analyze execution logs and calibrate task difficulty (Easy to Very Hard) using standard Terminal-Bench run configurations against various model tiers (Haiku, Sonnet, Opus).
  • Quality Assurance: Participate in a rigorous, 4-layer

EU Requirements

Job Details

Posted10 September 2026
Closes10 October 2026
Work ModeOnsite

Contact

Similar Jobs

Finding similar jobs...

AI Benchmark Engineer | Native Language Specialist - French (France) - Remote at LILT (Production) | EuroTalent AI