AI Quality & Evaluation Lead

Job not on LinkedIn

🔥 1 minute ago

🇵🇱 Poland – Remote

⏳ Contract/Temporary

🟠 Senior

🤖 Artificial Intelligence

👻 Ghost score 12%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Samsung Food

Samsung Food

51 - 200 employees

Founded 2012

👥 B2C

🍽️ Food & Beverage

☁️ SaaS

B2C • Food & Beverage • SaaS

Samsung Food is a consumer-facing app and web service for recipe saving, meal planning, grocery shopping, and recipe sharing. It offers a digital recipe box that saves recipes from any website, communities to discover and share recipes tailored to diets and preferences, a drag-and-drop weekly meal planner, automatic smart shopping list generation from recipes and meal plans, nutrition and calorie calculations for saved or custom recipes, and browser/mobile integrations (mobile apps, web app, Chrome extension). The platform emphasizes time-saving, health tracking, and social recipe discovery and is positioned as an award-nominated food/productivity app.

📋 Description

• Build the experience-quality evaluation layer for AI-generated coaching in Samsung Health • Author a failure taxonomy for coaching outputs • Hand-grade at least 150 traces with open-coded notes, including at least 40 sparse-data synthetic profiles • Categorize failures into 5–10 categories and quantify their frequency • Freeze at least 50 traces as a holdout set • Create binary pass/fail criteria for each taxonomy category • Write annotation guidelines with pass and fail examples • Double-code at least 30 traces, resolve disagreements, and maintain a guideline revision log • Create judge prompts for every criterion • Validate the LLM judge using true-positive and true-negative rates on development and holdout sets • Run weekly readouts on helpfulness, relevance, and tone • Document a revalidation routine triggered by model or prompt changes and quarterly regardless • Produce a complete playbook covering grading, taxonomy updates, rubric revisions, judge revalidation, and weekly readouts • Conduct a handover test enabling the internal owner to rerun validation and a weekly readout independently • Obtain Head of Product sign-off on the taxonomy and rubric • Collaborate with and hand over the method to an internal owner

🎯 Requirements

• Experience running the full evaluation loop at least once on conversational or generated-text output • Experience with error analysis on real traces • Experience building a failure taxonomy • Experience creating binary criteria and annotation guidelines • Experience validating an LLM judge against personal labels • Background in conversation design, AI product quality, model behaviour/policy, content design, human data operations, UX research with strong qualitative coding, RLHF, applied linguistics, or product management • Ability to identify previously unnamed failure modes by reading output • Ability to act as arbiter and document overrulings and rationale • Understanding that rubrics are discovered through grading • Ability to explain limitations of judge agreement rates • Ability to write unambiguous annotation guidelines • Comfort working in notebooks and spreadsheets • No production coding required • Nutrition, weight management, or behaviour change domain expertise is not required • Must be comfortable being responsible for own taxes and having no paid time off

🏖️ Benefits

• Remote-first work arrangement • Independent contractor setup • Flexible remote work • Global team of more than 100 people in over 30 countries

Apply Now

Similar Jobs

🕒 September 7

Welo Global

1001 - 5000

🤖 Artificial Intelligence

🤝 B2B

☁️ SaaS

Generative AI Analyst reviewing and annotating Polish content for Welo Data, a global AI data company. Improving datasets and quality processes for multilingual generative AI systems.

🗣️🇵🇱 Polish Required

🕒 July 24

CodiLime

201 - 500

🤝 B2B

📡 Telecommunications

🔧 Hardware

AI Validation Engineer at CodiLime developing Python-based frameworks for testing low-level SDK APIs. Engaging with clients and utilizing modern AI tools for code generation and validation.

🕒 May 28

10x.Team

11 - 50

💼 Consulting

📣 Marketing

📦 Logistics

Compliance Specialist reviewing and enhancing AI-generated compliance content in a remote role. Engage in a groundbreaking environment shaping AI compliance strategies while working flexible hours.

🇵🇱 Poland – Remote

💵 €90 - €180 / hour

💰 $1.1M Seed Round - 10x Team on 2024-05

⏳ Contract/Temporary

🟡 Mid-level

🟠 Senior

🤖 Artificial Intelligence

🕒 May 28

10x.Team

11 - 50

💼 Consulting

📣 Marketing

📦 Logistics

Surgeon needed to refine AI-generated surgical outputs at leading AI labs. Work remotely 8-20 hours weekly from EU or UK, enhancing AI relevance in medical systems.

🇵🇱 Poland – Remote

💵 €200 - €250 / hour

💰 $1.1M Seed Round - 10x Team on 2024-05

⏳ Contract/Temporary

🟡 Mid-level

🟠 Senior

🤖 Artificial Intelligence

🕒 May 28

10x.Team

11 - 50

💼 Consulting

📣 Marketing

📦 Logistics

Analytics Consultant improving AI-generated analytics outputs and real-world relevance. Freelance role requiring 8-20 hours a week, fully remote within the EU/UK.

🇵🇱 Poland – Remote

💵 €81 - €140 / hour

💰 $1.1M Seed Round - 10x Team on 2024-05

⏳ Contract/Temporary

🟡 Mid-level

🟠 Senior

🤖 Artificial Intelligence