AI Benchmark Engineer – Native Language Specialist, French

🔥 22 minutes ago

🇫🇷 France – Remote

⏳ Contract/Temporary

🟡 Mid-level

🟠 Senior

🤖 Artificial Intelligence

👻 Ghost score 25%

infoinfo

🗣️🇫🇷 French Required

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of LILT AI

LILT AI

201 - 500 employees

Founded 2015

🤖 Artificial Intelligence

☁️ SaaS

🏢 Enterprise

💰 $25M Funding Round - LILT on 2025-06

Artificial Intelligence • SaaS • Enterprise

LILT AI is a multilingual AI platform that helps enterprises and public-sector organizations create, translate, verify, and manage content across languages at scale. It combines domain-specific AI models, continuous model training, and a human intelligence layer of professional linguists to deliver secure, brand-consistent localization, translation verification, and multilingual model development. The platform offers enterprise project management, 100+ native integrations, on-prem and air-gapped deployments, and tools to train, evaluate, and govern multilingual AI for use cases like website localization, product launches, technical documentation, and regulatory/compliance communications.

📋 Description

• Design, build, and validate Terminal-Bench tasks for multilingual software challenges • Evaluate coding agents • Build realistic task environments using datasets and files in French, keeping assets in the target language • Find AI failure points through prompting and translation work in French • Support robust reference implementations • Write reliable, deterministic verifier scripts • Analyze execution logs and calibrate task difficulty from Easy to Very Hard • Run standard Terminal-Bench configurations against Haiku, Sonnet, and Opus model tiers • Participate in a four-layer human quality-control process: creation, human review, calibration review, and audit • Work with automated LLM-based checks to ensure fairness, grammatical accuracy, and benchmark integrity

🎯 Requirements

• 5+ years of industry experience in software engineering • Proven track record at leading technology companies and/or graduation from top-tier engineering universities • Native or near-native French fluency with deep understanding of grammar, register, and phrasing rules • High English proficiency • Strong proficiency in Python, standard shell scripting, and data processing • Extensive experience with Terminal/CLI-based development workflows • Working familiarity with coding agents • Deep technical understanding of multilingual text-processing pitfalls • Experience with encoding/decoding robustness and Unicode normalization • Knowledge of locale-dependent conventions, including collation, casing, and non-Gregorian dates • Knowledge of text I/O, toolchain interoperability, and safe string operations • For applicable languages, knowledge of bidirectional/RTL handling, font fallbacks, and rendering/typography • Updated CV in English • Successful completion of a GenAI assessment • Reliable availability and commitment; most tasks require at least 2 hours per day or 10 hours per week • Contractors must be able to comply with applicable geographic restrictions and are responsible for their own tax obligations

🏖️ Benefits

• Flexible schedule and ability to work as much or as little as desired • Competitive rates • Prompt payments • Access to diverse, innovative projects • Portfolio and skills development across industries and domains • Global community of linguists, subject matter experts, and language professionals • No fixed hours, check-ins, or micromanaging

Apply Now

Similar Jobs

🕒 August 12

Perle Systems

51 - 200

🏭 Manufacturing

📦 Logistics

💼 Consulting

Face-motion contributor recording short videos for Perle AI's computer vision training data. Completing simple camera-based tasks remotely across six eligible countries.

🕒 August 5

UNESCO

1001 - 5000

📚 Education

🤝 Non-profit

🌍 Social Impact

UNESCO Research Consultant comparing AI governance frameworks across seven jurisdictions. Producing policy profiles and recommendations for inclusive, ethical AI in cultural and creative sectors.

🕒 July 27

Welo Global

1001 - 5000

🤖 Artificial Intelligence

🤝 B2B

☁️ SaaS

Generative AI Analyst responsible for evaluating and annotating AI-generated content. Join Welo Data's global community and support next-generation Generative AI systems.

🗣️🇫🇷 French Required

🕒 July 22

Volga Partners

1001 - 5000

🤖 Artificial Intelligence

🤝 B2B

🏢 Enterprise

AI Evaluation & Annotation Reviewer evaluating AI-generated content in French for a major AI project. Engaging in linguistic analysis and feedback to improve AI system performance.

🗣️🇫🇷 French Required

🕒 June 9

RWS Group

5001 - 10000

💼 Consulting

🏥 Healthcare

⚖️ Legal

Speech AI Evaluation Specialist evaluating AI-generated content in French for flexibility and remote work. Ideal for students or gig workers seeking to shape AI development.

🗣️🇫🇷 French Required