AI Evaluation Guidelines & Rubric Specialist

🔥 36 minutes ago

🗽 New York – Remote

infoinfo

💵 $40 - $60 / hour

⏰ Full Time

🟡 Mid-level

🟠 Senior

🤖 Artificial Intelligence

👻 Ghost score 0%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of 24-MAG

24-MAG

2 - 10 employees

🤝 B2B

💼 Consulting

B2B • Consulting

24-MAG is a commercial strategy and execution firm that helps B2B organizations design and implement systems, workflows, and operating rhythms for sales, client management, and cross-functional projects. They focus on transforming scattered processes into aligned, measurable, and scalable commercial functions—covering pipeline structure, account management frameworks, and operational discipline for teams seeking efficient, intentional growth.

📋 Description

• Translate programme requirements into clear, structured, and actionable instructions for human evaluators • Develop guidelines applicable to standard scenarios and complex edge cases • Define terminology, rating criteria, decision rules, exceptions, and escalation pathways • Ensure instructions are accessible to raters while preserving domain-specific precision • Design detailed scoring rubrics for evaluating generative AI outputs • Establish measurable criteria for correctness, relevance, reasoning quality, completeness, and instruction adherence • Create examples and counterexamples illustrating different performance levels • Align evaluation frameworks with programme objectives and quality standards • Review draft specifications for ambiguity, contradiction, missing information, and inconsistent terminology • Identify instructions that may lead to conflicting interpretations across raters • Revise guideline sets for reliable application with minimal escalation • Document before-and-after improvements to written requirements and evaluation instructions • Convert specialist-domain specifications into rater-ready guidance across finance, retail, insurance, legal, sports, and other domains • Collaborate with subject matter experts on domain-specific terminology and professional judgment • Preserve technical nuance while making instructions clear to non-specialist evaluators • Maintain consistent structure and quality across domain-specific guideline sets • Participate in onboarding, calibration, documentation review, and ongoing guideline refinement

🎯 Requirements

• At least 3 years of professional experience in linguistics, instructional design, technical writing, content design, or a closely related field • Direct experience developing or refining guidelines and rubrics for human evaluators in generative AI, RLHF, or model-assessment programmes • Direct experience developing rater guidelines or rubrics for generative AI or RLHF programmes is required • Demonstrated ability to resolve ambiguity and contradiction in complex written specifications • Experience translating specialist requirements into clear and practical instructions • Ability to work effectively across multiple subject-matter domains • A portfolio or concrete examples showing measurable improvements to guidelines, rubrics, or instructional materials • Demonstrable professional growth and increasing responsibility • Reliable availability for at least 35 hours per week during weekdays • Equivalent professional experience in AI evaluation, technical documentation, or guideline development may also be considered • Applicants should be prepared to provide concrete examples of guideline or specification improvements • Immediate availability is preferred • Experience supporting large language model evaluation, reinforcement learning from human feedback, or AI training-data programmes • Familiarity with annotation platforms, human-feedback workflows, and rater calibration processes • Experience developing domain-specific guidance for finance, insurance, retail, legal, sports, or comparable fields • Knowledge of controlled language, information architecture, taxonomy design, or content governance • Experience conducting guideline usability tests or analysing inter-rater consistency • Familiarity with version control, documentation systems, and structured authoring tools • Previous collaboration with researchers, programme managers, engineers, and subject matter experts

🏖️ Benefits

• Full-time remote engagement • Competitive hourly compensation of $40–$60 per hour depending on expertise and project scope • W-2 contingent employment arrangement • Remote work from anywhere in the United States • Expected commitment of at least 35 hours per week during weekdays • Onboarding and calibration opportunities • Project scope and duration may be adjusted according to programme requirements and performance

Apply Now

Similar Jobs

🔥 12 hours ago

Workana

51 - 200

👥 HR Tech

🏪 Marketplace

🎯 Recruiter

AI web scraping engineer building stealth, self-healing scraping systems for Pinwheel API. Integrating anti-bot countermeasures, distributed infrastructure, and LLM parsing for US employment-data services.

🔥 12 hours ago

Tiger Analytics

1001 - 5000

🏥 Healthcare

📦 Logistics

📣 Marketing

AI Engagement Lead leading client engagements and hands-on GenAI engineering for Tiger Analytics, an advanced analytics consulting firm serving Fortune 100 companies.

🔥 14 hours ago

EXCEL Group

1001 - 5000

📦 Logistics

🏭 Manufacturing

💼 Consulting

Data & AI Analyst transforming insurance marketing data into actionable insights, dashboards, and AI-powered automation. Building scalable solutions with SQL, BigQuery, Tableau, Python, and emerging AI tools.

🔥 15 hours ago

Humana

10,000+ employees

🏥 Healthcare

🛡️ Insurance

⚕️ Healthcare Insurance

Lead architect shaping AI, cloud, data, and healthcare platforms for Humana’s CenterWell services. Governing secure, scalable solutions across regulated healthcare ecosystems.

🔥 19 hours ago

Palo Alto Networks

10,000+ employees

🔒 Cybersecurity

🏢 Enterprise

Palo Alto Networks cybersecurity leader building global distributor dashboards, data pipelines, and AI automations. Driving incentive validation and cloud marketplace integrations across channel operations.