AI Benchmark Engineer, Native Language Specialist – Chinese

🔥 4 minutes ago

🇹🇼 Taiwan – Remote

⏳ Contract/Temporary

🟡 Mid-level

🟠 Senior

🤖 Artificial Intelligence

👻 Ghost score 25%

infoinfo

🗣️🇨🇳 Chinese Required

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of LILT AI

LILT AI

201 - 500 employees

Founded 2015

🤖 Artificial Intelligence

☁️ SaaS

🏢 Enterprise

💰 $25M Funding Round - LILT on 2025-06

Artificial Intelligence • SaaS • Enterprise

LILT AI is a multilingual AI platform that helps enterprises and public-sector organizations create, translate, verify, and manage content across languages at scale. It combines domain-specific AI models, continuous model training, and a human intelligence layer of professional linguists to deliver secure, brand-consistent localization, translation verification, and multilingual model development. The platform offers enterprise project management, 100+ native integrations, on-prem and air-gapped deployments, and tools to train, evaluate, and govern multilingual AI for use cases like website localization, product launches, technical documentation, and regulatory/compliance communications.

📋 Description

• Design, build, and validate Terminal-Bench tasks for multilingual software challenges • Evaluate coding agents • Build realistic task environments using datasets and files in Chinese (Taiwan) • Find AI failure points through prompting and translation in the native language • Develop robust reference implementations • Write reliable, deterministic verifier scripts • Analyze execution logs and calibrate task difficulty from Easy to Very Hard • Run standard Terminal-Bench configurations against Haiku, Sonnet, and Opus model tiers • Participate in a four-layer human quality-control process covering creation, human review, calibration review, and audit • Work alongside automated LLM-based checks to ensure fairness, grammatical accuracy, and benchmark integrity

🎯 Requirements

• 5+ years of industry experience in software engineering • Proven track record at leading technology companies and/or graduation from top-tier engineering universities • Native or near-native fluency in Chinese (Taiwan), with deep understanding of grammar, register, and phrasing rules • High English proficiency • Strong proficiency in Python • Strong proficiency in standard shell scripting • Strong proficiency in data processing • Extensive experience with Terminal/CLI-based development workflows • Working familiarity with coding agents • Deep technical understanding of multilingual text processing pitfalls • Experience with encoding/decoding robustness and Unicode normalization • Experience with locale-dependent conventions, including collation, casing, and non-Gregorian dates • Experience with text I/O, toolchain interoperability, and safe string operations • Updated CV in English • Successful completion of a GenAI assessment • Reliable availability and commitment; most tasks require a minimum of 2 hours per day or 15–20 hours per week • Contractors must not be located in regions subject to international embargoes or sanctions • Contractors are responsible for their own tax obligations

🏖️ Benefits

• Flexible schedule with no fixed hours, check-ins, or micromanaging • Competitive rates • Prompt payments • Access to diverse, innovative projects • Portfolio and skills development across industries and domains • Global community of linguists, subject matter experts, and language professionals • No health insurance, paid time off, or retirement contributions are provided • Hours are not guaranteed • Supplemental-income opportunity for independent 1099 contractors

Apply Now

Similar Jobs

🕒 August 21

Welo Global

1001 - 5000

🤖 Artificial Intelligence

🤝 B2B

☁️ SaaS

Generative AI Senior Linguist reviewing Simplified and Traditional Chinese annotations for Welo Data’s AI quality projects. Ensuring linguistic accuracy, consistency, and guideline compliance.

🗣️🇨🇳 Chinese Required

🕒 August 5

Welo Global

1001 - 5000

🤖 Artificial Intelligence

🤝 B2B

☁️ SaaS

Generative AI Analyst reviewing Traditional Chinese text, images, audio, and video for Welo Data’s AI training projects. Ensuring quality and improving annotation guidelines.

🗣️🇨🇳 Chinese Required

🕒 July 31

Welo Global

1001 - 5000

🤖 Artificial Intelligence

🤝 B2B

☁️ SaaS

Generative AI Analyst contributing to next-generation AI systems at Welo Data. Involves content evaluation, annotation, and quality checks across various media formats.

🗣️🇨🇳 Chinese Required