AI Evals Engineer – Evaluation Datasets, Ground Truth

🕒 vor 25 Tagen

🇺🇸 Vereinigte Staaten – Remote

⏰ Vollzeit

🟡 Mittelstufe

🟠 Senior

🤖 Künstliche Intelligenz

👻 Geisterscore 15%

infoinfo

🗣️🇺🇸🇬🇧 Englisch erforderlich

Jetzt Bewerben
Ähnliche Remote-Jobs finden

📊 Überprüfen Sie Ihre Lebenslauf-Bewertung für diese Stelle

Verbessern Sie Ihre Chancen auf ein Vorstellungsgespräch, indem Sie Ihre Lebenslauf-Bewertung vor der Bewerbung überprüfen.

Logo of Prophetic

Prophetic

11 - 50 Mitarbeiter

🏠 Immobilien

☁️ SaaS

🤖 Künstliche Intelligenz

💰 €3.200.000 Seed im 2025-05

Real Estate • SaaS • Artificial Intelligence

Prophetic ist eine KI-native SaaS-Plattform, die Land- und Entwicklungsinformationen für Immobilienentwickler und verwandte Fachleute bereitstellt. Die Plattform entdeckt marktferne Parzellen und Assemblierungen, liefert sofortige Klarheit über die Zoneneinteilung, führt Machbarkeits-/Ertragsstudien in Sekunden durch, verfolgt Entwicklungsaktivitäten und zentralisiert Landpipeline- und Sorgfaltsprüfung-Daten. Zu den Hauptprodukten gehören SearchAI, ZoneAI, SiteAI, DevMap, ein Land Relationship Manager (LRM) und Due-Diligence-Tools, die darauf ausgelegt sind, Deals zu beschleunigen, manuelle Recherchen zu reduzieren und Entwicklern zu helfen, schneller Entscheidungen zu treffen und mehr Chancen zu gewinnen.

Beschreibung

• Decompose pipelines and system architecture into evaluable modules with explicit input-to-expected-output contracts • Define correctness through rubrics, label schemas, edge-case policies, and business-relevant tolerances • Prioritize validation sets that unlock the most iteration • Pull stratified, de-identified production samples reflecting real traffic and long-tail cases • Scope and run human-labeling programs, create annotation guidelines and calibration sets, measure inter-annotator agreement, and manage vendors or subject-matter experts • Design synthetic/oracle-generated ground truth and verify oracle outputs against human-labeled samples • Create programmatic and adversarial labels, templated edge cases, and incident-based backtests • Version datasets and track lineage, splits, model/prompt exposure, contamination, and leakage • Slice datasets by customer segment, input type, and difficulty; refresh and retire examples as traffic changes • Calibrate automated graders against human gold sets and determine when they are reliable enough for gating • Report precision/recall/F1, confusion matrices, calibration, confidence intervals, and sample-size requirements • Partner with eval-harness and CI engineers and ML engineers on classifier features • Report to the Chief AI Officer and work across product and pipeline teams; expected to eventually lead a small eval-engineering team

🎯 Anforderungen

• Candidates must be authorized to work in the United States without current or future sponsorship • Experience building or evaluating ML or LLM-powered systems in production, in a role where output quality was your problem • Experience building evaluation or validation datasets, with ability to explain correctness definitions, label sourcing, failure modes, and label quality • Working fluency in ML validation fundamentals: train/validation/test discipline, stratified sampling, precision/recall trade-offs, class imbalance, calibration, inter-rater agreement, basic significance testing and power • Understanding of feature engineering sufficient to reason about classifier failures and data requirements • Strong Python and SQL; comfortable pulling and reshaping data • Hands-on experience with LLM-based systems, including prompting, structured outputs, agent/tool-use harnesses, non-determinism, prompt sensitivity, and evaluator bias • Judgment about when LLM-as-judge is reliable and how to validate it • Ability to read system designs, understand business logic, and translate it into label schemas • Experience running human-labeling programs end to end, including vendor selection, guideline authoring, QA sampling, and cost/quality trade-offs (nice to have) • Experience with eval tooling (nice to have) • Experience with data labeling platforms (nice to have) • Experience using a frontier model as a distillation/oracle source (nice to have)

🏖️ Vorteile

• 100% medical, dental & vision insurance coverage for you; 30% coverage for dependents • Competitive salary and meaningful early-stage equity • Unlimited PTO • Hybrid/Remote stipend • In-office perks: snacks, drinks, coffee, ping-pong table, and more • Budget for intra-office travel • 2–3 annual team meetups in person • Generous budget for labeling vendors and oracle compute

Jetzt Bewerben

Ähnliche Jobs

🕒 vor 25 Tagen

Juniper Square

201 - 500

💸 Finanzen

🏠 Immobilien

☁️ SaaS

Forward Deployed Engineer building custom AI agents for Juniper Square’s private-markets operations platform. Owning client discovery, full-stack delivery, deployment, and adoption.

🇺🇸 Vereinigte Staaten – Remote

💵 $165.000 - $200.000 / Jahr

💰 €75.000.000 Series C im 2019-11

⏰ Vollzeit

🟠 Senior

🤖 Künstliche Intelligenz

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 26 Tagen

EQL Tech (sales & engineering talent)

1 - 10

🎯 Rekrutierung

🤖 Künstliche Intelligenz

🤝 B2B

AI Resilience Partner building Focused Research Organizations for Convergent Research, a nonprofit science studio. Developing portfolios, funding projects, recruiting founders, and guiding boards.

🇺🇸 Vereinigte Staaten – Remote

💵 $200.000 - $350.000 / Jahr

⏰ Vollzeit

🟡 Mittelstufe

🟠 Senior

🤖 Künstliche Intelligenz

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 26 Tagen

Rockstar

1 - 10

💼 Beratung

📣 Marketing

📦 Logistik

AI transformation specialist helping mid-market companies turn AI investment into measurable business value. Leading executive sessions, behavior change, client adoption, and relationship management.

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 26 Tagen

Clinton Health Access Initiative, Inc.

1001 - 5000

🏥 Gesundheitswesen

💼 Beratung

📦 Logistik

Market-shaping manager expanding AI diagnostic access across low- and middle-income countries for CHAI. Leading supplier negotiations, policy, financing, and country scale-up strategies.

🇺🇸 Vereinigte Staaten – Remote

⏰ Vollzeit

🟠 Senior

🔴 Experte

🤖 Künstliche Intelligenz

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 27 Tagen

Pragmatico

2 - 10

💼 Beratung

🤖 Künstliche Intelligenz

🤝 B2B

AI transformation facilitator helping mid-market companies turn investment into measurable adoption and business value. Coaching executives, leading practical sessions, and sustaining client engagement.

🇺🇸 Vereinigte Staaten – Remote

⏰ Vollzeit

🟡 Mittelstufe

🟠 Senior

🤖 Künstliche Intelligenz

🗣️🇺🇸🇬🇧 Englisch erforderlich