AI Safety Expert, English – Dutch

🕒 August 10

🇺🇸 United States – Remote

💵 $48 - $62 / hour

⏳ Contract/Temporary

🟡 Mid-level

🟠 Senior

🤖 Artificial Intelligence

👻 Ghost score 12%

infoinfo

🗣️🇳🇱 Dutch Required

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Mercor

Mercor

51 - 200 employees

Founded 2023

🔥 Funding within the last year

💰 $350M Series C - Mercor on 2025-10

Mercor is a company for which no descriptive text was provided in the input. Additional information (products, services, target customers, or industry specifics) is needed to create an accurate summary and select appropriate industries.

📋 Description

• Red team conversational AI models and agents using jailbreaks, prompt injections, misuse cases, bias exploitation, and multi-turn manipulation • Generate human data by annotating failures, classifying vulnerabilities, and flagging systemic risks • Apply taxonomies, benchmarks, and playbooks to keep testing consistent • Produce reproducible reports, datasets, and attack cases for customers • Probe AI outputs involving sensitive topics such as bias, misinformation, and harmful behaviors • Uncover vulnerabilities that automated tests miss • Expand evaluation coverage across more scenarios • Strengthen customer AI systems and improve their safety, robustness, and trustworthiness • Collaborate with leading researchers on projects training and enhancing AI systems

🎯 Requirements

• Prior red teaming experience in AI adversarial work, cybersecurity, or socio-technical probing • Native fluency in English and Dutch • Experience with adversarial inputs and probing AI models for vulnerabilities • Ability to use frameworks, taxonomies, benchmarks, or playbooks for structured testing • Ability to explain risks clearly to technical and non-technical stakeholders • Ability to adapt across projects and customers • H1-B and STEM OPT candidates are not supported • Nice-to-have: adversarial ML experience, including jailbreak datasets, prompt injection, RLHF/DPO attacks, or model extraction • Nice-to-have: cybersecurity experience, including penetration testing, exploit development, or reverse engineering • Nice-to-have: socio-technical risk experience, including harassment/disinformation probing, abuse analysis, or conversational AI testing • Nice-to-have: creative probing experience in psychology, acting, or writing

🏖️ Benefits

• Fully remote work • Flexible schedule / work on your own schedule • Weekly payments via Stripe or Wise • Higher-sensitivity project participation is optional • Clear guidelines and wellness resources for higher-sensitivity projects • Competitive pay • Collaboration with leading researchers • Referral bonus of up to $250 per successful referral

Apply Now

Similar Jobs

🕒 August 7

Lightly

11 - 50

🤖 Artificial Intelligence

🤝 B2B

☁️ SaaS

Energy sector expert evaluating AI-generated forecasts for Lightly AG, an ETH and HSG spin-off building machine-learning and computer-vision technology. Assessing forecast accuracy and response quality using a provided rubric.

🕒 August 6

Cayuse Holdings

501 - 1000

💼 Consulting

📦 Logistics

🏥 Healthcare

AI Native Engineer building Claude Code agents, skills, and MCP integrations. Developing AWS-backed harness components for retrieval, tools, evaluation, and observability.

🕒 August 6

Terac

1 - 10

🤖 Artificial Intelligence

🤝 B2B

STEM experts creating graduate-to-PhD-level problems and exact solutions for Terac’s AI training research. Reviewing AI responses and identifying scientific reasoning errors.

🕒 August 6

Hippocratic AI

11 - 50

🤖 Artificial Intelligence

⚕️ Healthcare Insurance

🏥 Healthcare

Evaluating Hippocratic AI’s clinical agents for symptom triage, diabetes medication titration, and travel health guidance. Providing safety-focused feedback to improve AI-driven healthcare tools.

🕒 August 5

DevFixr

2 - 10

🎯 Recruiter

🤝 B2B

💼 Consulting

AI Trainer creating realistic, unambiguous software engineering tasks for frontier language model training. Evaluating model performance across Python and other programming-language environments.