AI Evals Engineer – Evaluation Datasets, Ground Truth

Job not on LinkedIn

🔥 21 hours ago

🇺🇸 United States – Remote

⏰ Full Time

🟡 Mid-level

🟠 Senior

🤖 Artificial Intelligence

👻 Ghost score 17%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Prophetic

Prophetic

11 - 50 employees

🏠 Real Estate

☁️ SaaS

🤖 Artificial Intelligence

💰 $3.2M Seed on 2025-05

Real Estate • SaaS • Artificial Intelligence

Prophetic is an AI-native SaaS platform that provides land and development intelligence for real estate developers and related professionals. The platform uncovers off-market parcels and assemblies, delivers instant zoning clarity, performs feasibility/yield studies in seconds, tracks development activity, and centralizes land pipeline and due-diligence data. Key products include SearchAI, ZoneAI, SiteAI, DevMap, a Land Relationship Manager (LRM), and Due Diligence tools designed to accelerate deals, reduce manual research, and help developers decide faster and win more opportunities.

📋 Description

• Decompose pipelines and system architecture into evaluable modules with explicit input-to-expected-output contracts • Define correctness through rubrics, label schemas, edge-case policies, and business-relevant tolerances • Prioritize validation sets that unlock the most iteration • Pull stratified, de-identified production samples reflecting real traffic and long-tail cases • Scope and run human-labeling programs, create annotation guidelines and calibration sets, measure inter-annotator agreement, and manage vendors or subject-matter experts • Design synthetic/oracle-generated ground truth and verify oracle outputs against human-labeled samples • Create programmatic and adversarial labels, templated edge cases, and incident-based backtests • Version datasets and track lineage, splits, model/prompt exposure, contamination, and leakage • Slice datasets by customer segment, input type, and difficulty; refresh and retire examples as traffic changes • Calibrate automated graders against human gold sets and determine when they are reliable enough for gating • Report precision/recall/F1, confusion matrices, calibration, confidence intervals, and sample-size requirements • Partner with eval-harness and CI engineers and ML engineers on classifier features • Report to the Chief AI Officer and work across product and pipeline teams; expected to eventually lead a small eval-engineering team

🎯 Requirements

• Candidates must be authorized to work in the United States without current or future sponsorship • Experience building or evaluating ML or LLM-powered systems in production, in a role where output quality was your problem • Experience building evaluation or validation datasets, with ability to explain correctness definitions, label sourcing, failure modes, and label quality • Working fluency in ML validation fundamentals: train/validation/test discipline, stratified sampling, precision/recall trade-offs, class imbalance, calibration, inter-rater agreement, basic significance testing and power • Understanding of feature engineering sufficient to reason about classifier failures and data requirements • Strong Python and SQL; comfortable pulling and reshaping data • Hands-on experience with LLM-based systems, including prompting, structured outputs, agent/tool-use harnesses, non-determinism, prompt sensitivity, and evaluator bias • Judgment about when LLM-as-judge is reliable and how to validate it • Ability to read system designs, understand business logic, and translate it into label schemas • Experience running human-labeling programs end to end, including vendor selection, guideline authoring, QA sampling, and cost/quality trade-offs (nice to have) • Experience with eval tooling (nice to have) • Experience with data labeling platforms (nice to have) • Experience using a frontier model as a distillation/oracle source (nice to have)

🏖️ Benefits

• 100% medical, dental & vision insurance coverage for you; 30% coverage for dependents • Competitive salary and meaningful early-stage equity • Unlimited PTO • Hybrid/Remote stipend • In-office perks: snacks, drinks, coffee, ping-pong table, and more • Budget for intra-office travel • 2–3 annual team meetups in person • Generous budget for labeling vendors and oracle compute

Apply Now

Similar Jobs

🔥 21 hours ago

Palo Alto Networks

10,000+ employees

🔒 Cybersecurity

🏢 Enterprise

Analytics manager architecting dashboards, distributor data pipelines, incentive validation, and AI automation for Palo Alto Networks’ global cybersecurity channel. Integrating cloud marketplaces and operational intelligence.

🕒 Yesterday

The Cigna Group

10,000+ employees

🏥 Healthcare

🛡️ Insurance

AI strategy leader prioritizing high-value investments for The Cigna Group’s healthcare businesses. Framing opportunities, validating value, and guiding disciplined funding decisions.

🕒 Yesterday

High Bridge Consulting LLC

51 - 200

💼 Consulting

📣 Marketing

📦 Logistics

Senior AI engineer designing and prototyping Agentic AI solutions for High Bridge, a technology consulting firm. Translating stakeholder problems into multi-agent workflows and guiding implementation.

🕒 Yesterday

GE Vernova

10,000+ employees

💼 Consulting

📦 Logistics

🏭 Manufacturing

AI delivery lead advancing analytics and AI products for GE Vernova’s Global Services & Commercial organization. Owning strategy, transformation, governance, and production delivery.

🕒 Yesterday

GE Vernova

10,000+ employees

💼 Consulting

📦 Logistics

🏭 Manufacturing

AI delivery leader advancing GE Vernova’s Global Services & Commercial analytics and AI products. Driving data readiness, transformation initiatives, production deployment, and measurable business value.