Senior Evaluation Algorithm Engineer

Job not on LinkedIn

🔥 6 hours ago

🌐 Hong Kong, Taiwan, +1 more countries – Remote

infoinfo

⏰ Full Time

🟠 Senior

👷🏻‍♀️ Engineer

👻 Ghost score 11%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Binance

Binance

1001 - 5000 employees

Founded 2017

₿ Crypto

💳 Fintech

💰 Initial Coin Offering on 2020-12

Crypto • Fintech

Binance is the world's leading cryptocurrency exchange, serving over 235 million registered users across more than 180 countries. The platform offers a wide array of services, including the trading of over 350 cryptocurrencies in Spot, Margin, and Futures markets. Users can also buy and sell crypto via Binance P2P, earn interest through Binance Earn, and engage in NFT trading on the Binance NFT marketplace. Binance provides low transaction fees and diverse payment options, making it a preferred choice for cryptocurrency enthusiasts worldwide.

📋 Description

• Design end-to-end LLM evaluation plans for dialogue, financial trading, and other business scenarios • Build evaluation metric systems and rubrics that produce quantifiable, reproducible, and explainable conclusions • Lead evaluation dataset design and construction • Define evaluation dimensions and scenario coverage • Establish data annotation guidelines and quality-control processes • Build benchmarks reflecting business needs with discriminative power • Analyze model capability boundaries and failure modes • Produce actionable improvement recommendations and collaborate with algorithm and product teams to drive model iteration • Automate and scale evaluation workflows • Build sustainable evaluation platforms and toolchains • Collaborate with algorithm, product, and data teams to translate business and model objectives into evaluation standards • Turn evaluation findings into concrete R&D directions and drive implementation

🎯 Requirements

• Master's degree or above in Computer Science, Artificial Intelligence, Mathematics, Statistics, or related fields • Solid algorithmic foundation and understanding of LLM principles, training, and fine-tuning processes • Hands-on LLM evaluation experience at a large tech company • Participation in commercial deployment evaluation, not purely academic or offline benchmarking • Familiarity with the full pipeline from evaluation data preparation and rubric design to evaluation-driven R&D • Familiarity with human evaluation, model-based automatic evaluation / LLM-as-a-judge, and metric computation • Ability to define appropriate evaluation dimensions for different business scenarios • Ability to write clear, actionable, and discriminative rubrics • Systematic control over evaluation data representativeness, annotation consistency, and result reliability • Proficient in Python • Experience in evaluation workflow automation, benchmark construction, or evaluation platform development • Ability to independently handle data processing, evaluation script writing, and result analysis • Strong business understanding and communication skills • Ability to translate evaluation findings into clear improvement directions and drive cross-team collaboration • Experience evaluating dialogue systems, AI Agents, or financial/trading LLMs (bonus) • Experience building high-quality AI training/evaluation data or data annotation systems (bonus) • Familiarity with RLHF, reward models, or preference data-related work (bonus)

🏖️ Benefits

• Competitive salary and company benefits • Work-from-home arrangement (the arrangement may vary depending on the work nature of the business team) • Opportunities for career growth and continuous learning • Collaborate with world-class talent in a user-centric global organization with a flat structure • Tackle unique, fast-paced projects with autonomy in an innovative environment • Results-driven workplace

Apply Now