AI Quality & Evaluation Lead

Stelle nicht auf LinkedIn

🕒 vor 2 Tagen

🇵🇱 Polen – Remote

⏳ Vertrag

🟠 Senior

🤖 Künstliche Intelligenz

👻 Geisterscore 12%

infoinfo

🗣️🇺🇸🇬🇧 Englisch erforderlich

Jetzt Bewerben
Ähnliche Remote-Jobs finden

📊 Überprüfen Sie Ihre Lebenslauf-Bewertung für diese Stelle

Verbessern Sie Ihre Chancen auf ein Vorstellungsgespräch, indem Sie Ihre Lebenslauf-Bewertung vor der Bewerbung überprüfen.

Logo of Samsung Food

Samsung Food

51 - 200 Mitarbeiter

Gegründet 2012

👥 B2C

🍽️ Lebensmittel & Getränke

☁️ SaaS

B2C • Food & Beverage • SaaS

Samsung Food ist eine verbraucherorientierte App und Web-Dienstleistung für das Speichern von Rezepten, die Essensplanung, den Lebensmitteleinkauf und das Teilen von Rezepten. Es bietet eine digitale Rezeptsammlung, die Rezepte von jeder Website speichert, Communities, um Rezepte zu entdecken und zu teilen, die auf Diäten und Vorlieben zugeschnitten sind, einen Drag-and-Drop-Wochenmahlzeitenplaner, eine automatische smarte Einkaufsliste, die aus Rezepten und Mahlzeitenplänen generiert wird, Nährwert- und Kalorienberechnungen für gespeicherte oder benutzerdefinierte Rezepte sowie Browser- und Mobilintegrationen (Mobile Apps, Web-App, Chrome-Erweiterung). Die Plattform hebt Zeitersparnis, Gesundheitsverfolgung und das Entdecken von sozialen Rezepten hervor und positioniert sich als preisgekrönte Lebensmittel-/Produktivitäts-App.

Beschreibung

• Build the experience-quality evaluation layer for AI-generated coaching in Samsung Health • Author a failure taxonomy for coaching outputs • Hand-grade at least 150 traces with open-coded notes, including at least 40 sparse-data synthetic profiles • Categorize failures into 5–10 categories and quantify their frequency • Freeze at least 50 traces as a holdout set • Create binary pass/fail criteria for each taxonomy category • Write annotation guidelines with pass and fail examples • Double-code at least 30 traces, resolve disagreements, and maintain a guideline revision log • Create judge prompts for every criterion • Validate the LLM judge using true-positive and true-negative rates on development and holdout sets • Run weekly readouts on helpfulness, relevance, and tone • Document a revalidation routine triggered by model or prompt changes and quarterly regardless • Produce a complete playbook covering grading, taxonomy updates, rubric revisions, judge revalidation, and weekly readouts • Conduct a handover test enabling the internal owner to rerun validation and a weekly readout independently • Obtain Head of Product sign-off on the taxonomy and rubric • Collaborate with and hand over the method to an internal owner

🎯 Anforderungen

• Experience running the full evaluation loop at least once on conversational or generated-text output • Experience with error analysis on real traces • Experience building a failure taxonomy • Experience creating binary criteria and annotation guidelines • Experience validating an LLM judge against personal labels • Background in conversation design, AI product quality, model behaviour/policy, content design, human data operations, UX research with strong qualitative coding, RLHF, applied linguistics, or product management • Ability to identify previously unnamed failure modes by reading output • Ability to act as arbiter and document overrulings and rationale • Understanding that rubrics are discovered through grading • Ability to explain limitations of judge agreement rates • Ability to write unambiguous annotation guidelines • Comfort working in notebooks and spreadsheets • No production coding required • Nutrition, weight management, or behaviour change domain expertise is not required • Must be comfortable being responsible for own taxes and having no paid time off

🏖️ Vorteile

• Remote-first work arrangement • Independent contractor setup • Flexible remote work • Global team of more than 100 people in over 30 countries

Jetzt Bewerben

Ähnliche Jobs

🕒 vor 1 Monat

Welo Global

1001 - 5000

🤖 Künstliche Intelligenz

🤝 B2B

☁️ SaaS

Generative AI Analyst reviewing and annotating Polish content for Welo Data, a global AI data company. Improving datasets and quality processes for multilingual generative AI systems.

🇵🇱 Polen – Remote

💵 $10 / Stunde

⏳ Vertrag

🟡 Mittelstufe

🟠 Senior

🤖 Künstliche Intelligenz

🗣️🇵🇱 Polnisch erforderlich

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 2 Monaten

CodiLime

201 - 500

🤝 B2B

📡 Telekommunikation

🔧 Hardware

AI Validation Engineer at CodiLime developing Python-based frameworks for testing low-level SDK APIs. Engaging with clients and utilizing modern AI tools for code generation and validation.

🇵🇱 Polen – Remote

💵 zł18.000 - zł22.000 / Monat

⏳ Vertrag

🟡 Mittelstufe

🟠 Senior

🤖 Künstliche Intelligenz

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 4 Monaten

10x.Team

11 - 50

💼 Beratung

📣 Marketing

📦 Logistik

Compliance Specialist reviewing and enhancing AI-generated compliance content in a remote role. Engage in a groundbreaking environment shaping AI compliance strategies while working flexible hours.

🇵🇱 Polen – Remote

💵 €90 - €180 / Stunde

💰 €1.085.129 Seed Round - 10x Team im 2024-05

⏳ Vertrag

🟡 Mittelstufe

🟠 Senior

🤖 Künstliche Intelligenz

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 4 Monaten

10x.Team

11 - 50

💼 Beratung

📣 Marketing

📦 Logistik

Surgeon needed to refine AI-generated surgical outputs at leading AI labs. Work remotely 8-20 hours weekly from EU or UK, enhancing AI relevance in medical systems.

🇵🇱 Polen – Remote

💵 €200 - €250 / Stunde

💰 €1.085.129 Seed Round - 10x Team im 2024-05

⏳ Vertrag

🟡 Mittelstufe

🟠 Senior

🤖 Künstliche Intelligenz

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 4 Monaten

10x.Team

11 - 50

💼 Beratung

📣 Marketing

📦 Logistik

Analytics Consultant improving AI-generated analytics outputs and real-world relevance. Freelance role requiring 8-20 hours a week, fully remote within the EU/UK.

🇵🇱 Polen – Remote

💵 €81 - €140 / Stunde

💰 €1.085.129 Seed Round - 10x Team im 2024-05

⏳ Vertrag

🟡 Mittelstufe

🟠 Senior

🤖 Künstliche Intelligenz

🗣️🇺🇸🇬🇧 Englisch erforderlich