AI Evaluation Specialist

🔥 15 hours ago

🌐 United States, Canada, +4 more countries – Remote

infoinfo

🗽 New York – Remote

infoinfo

💵 $30 - $90 / hour

⏱ Part Time

🟡 Mid-level

🟠 Senior

🤖 Artificial Intelligence

👻 Ghost score 0%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of 24-MAG

24-MAG

2 - 10 employees

🤝 B2B

💼 Consulting

B2B • Consulting

24-MAG is a commercial strategy and execution firm that helps B2B organizations design and implement systems, workflows, and operating rhythms for sales, client management, and cross-functional projects. They focus on transforming scattered processes into aligned, measurable, and scalable commercial functions—covering pipeline structure, account management frameworks, and operational discipline for teams seeking efficient, intentional growth.

📋 Description

• Evaluate AI-generated outputs against detailed rubrics, guidelines, and defined quality standards • Assess responses for accuracy, relevance, completeness, reasoning quality, and adherence to instructions • Apply consistent and impartial judgement across large volumes of evaluation examples • Identify outputs that fail to satisfy important quality or task requirements • Maintain reliable assessment standards across repeated evaluation workflows • Identify reasoning gaps, logic errors, inconsistencies, unsupported conclusions, and tool-use failures • Analyse where AI-generated outputs diverge from expected reasoning or quality standards • Document recurring model weaknesses and opportunities for improvement • Produce clear, concise, and actionable written feedback on strengths and areas for improvement • Explain the reasoning behind evaluation decisions and quality scores • Maintain detailed, transparent, traceable, and reproducible assessment documentation • Participate in rubric interpretation and ambiguous-case discussions • Help refine assessment criteria as AI models and project requirements develop • Contribute insights supporting process optimisation and evaluation best practices • Collaborate with other reviewers to improve alignment and reliability across evaluation workflows

🎯 Requirements

• Experience in grading, quality assurance, editorial review, assessment, annotation, or another field requiring careful analysis and detailed feedback • Advanced, regular use of AI assistants such as ChatGPT, Claude, or comparable tools for professional work and productivity • Strong ability to synthesise complex information and communicate conclusions clearly in writing • Experience with process improvement, rubric development, operational quality assessment, or structured evaluation workflows is advantageous • Strong critical-thinking skills with particular emphasis on consistency, integrity, and fairness • High attention to detail and comfort reviewing large volumes of similar examples • Ability to work independently while maintaining consistent evaluation quality • Collaborative approach to discussing ambiguous cases and refining shared assessment standards • Excellent written English and professional documentation skills • Based in the United States, Canada, United Kingdom, Ireland, Australia, or New Zealand • Authorised to undertake contract work in the relevant country • No prior formal experience in AI research or model training is required

🏖️ Benefits

• Part-time independent contractor engagement • Fully remote work • Flexible project scope, workload, timing, and duration depending on project requirements

Apply Now

Similar Jobs

🔥 16 hours ago

Calbright College

11 - 50

💼 Consulting

🏥 Healthcare

📦 Logistics

Adjunct Generative AI Instructor teaching adult learners through competency-based online education. Providing AI expertise, personalized support, feedback, and mentorship at Calbright College.

🕒 2 days ago

Full Stack Academy

11 - 50

📚 Education

Part-time online instructor teaching Agentic AI, LLMs, RAG, and automation. Simplilearn delivers immersive technology bootcamp training to professionals worldwide.

🕒 3 days ago

Full Stack Academy

11 - 50

📚 Education

Part-time Agentic AI instructor teaching live online classes for Simplilearn’s technology bootcamp programs. Mentoring adult learners, grading assignments, and connecting AI concepts to industry applications.

🕒 4 days ago

American Military University

1001 - 5000

💼 Consulting

🎖️ Defense

🏥 Healthcare

Part-time online AI faculty teaching undergraduate and graduate students for American Public University System. Delivering lessons, moderating forums, grading work, and supporting student success.

🕒 September 4

Anderson University (SC)

501 - 1000

📚 Education

Part-time doctoral faculty teaching AI, ethics, cybersecurity, and digital transformation in Anderson University’s professional Ed.D. program. Mentoring applied research and capstone projects for experienced professionals.