
2 - 10 employees
🤝 B2B
💼 Consulting
B2B • Consulting
24-MAG is a commercial strategy and execution firm that helps B2B organizations design and implement systems, workflows, and operating rhythms for sales, client management, and cross-functional projects. They focus on transforming scattered processes into aligned, measurable, and scalable commercial functions—covering pipeline structure, account management frameworks, and operational discipline for teams seeking efficient, intentional growth.
🔥 0 minutes ago
🗽 New York – Remote
💵 $45 - $60 / hour
⏳ Contract/Temporary
🟡 Mid-level
🟠 Senior
🤖 Artificial Intelligence
🗣️🇫🇮 Finnish Required
Improve your chances of getting an interview by checking your resume score before you apply.

2 - 10 employees
🤝 B2B
💼 Consulting
B2B • Consulting
24-MAG is a commercial strategy and execution firm that helps B2B organizations design and implement systems, workflows, and operating rhythms for sales, client management, and cross-functional projects. They focus on transforming scattered processes into aligned, measurable, and scalable commercial functions—covering pipeline structure, account management frameworks, and operational discipline for teams seeking efficient, intentional growth.
• Conduct adversarial testing of conversational AI models and agents • Develop jailbreaks, prompt-injection scenarios, misuse cases, and multi-turn manipulation strategies • Probe models for vulnerabilities across diverse conversational and adversarial scenarios • Apply systematic testing frameworks and established evaluation methodologies • Identify and classify model failures and safety vulnerabilities • Evaluate bias, misinformation, misuse, and potentially harmful model behaviours • Identify recurring or systemic patterns across model responses • Assess vulnerability severity, reproducibility, and practical significance • Produce high-quality human evaluation data from red teaming activities • Annotate model failures and categorise vulnerabilities • Create structured attack cases and supporting evaluation materials • Maintain consistency across repeated assessments and datasets • Document adversarial scenarios clearly and reproducibly • Produce reports, datasets, and structured findings for technical teams • Explain identified risks to technical and non-technical stakeholders • Record testing methodology, model behaviour, and relevant failure patterns • Contribute to broader evaluation coverage across models and use cases
• Native-level fluency in both English and Finnish is required • Prior experience with AI red teaming, adversarial AI evaluation, cybersecurity, or socio-technical system testing • Strong understanding of conversational AI systems and model failure modes • Experience developing structured adversarial tests and evaluation frameworks • Ability to identify subtle vulnerabilities and recurring behavioural patterns • Strong analytical reasoning and written communication skills • Ability to document findings clearly and reproducibly • Comfort working across changing projects, scenarios, and evaluation frameworks • Educational background in computer science, cybersecurity, artificial intelligence, machine learning, linguistics, behavioural science, or a related discipline may be helpful • Equivalent professional experience in AI safety, adversarial testing, security, or structured model evaluation may be considered • Practical red teaming experience is particularly valuable • Familiarity with adversarial machine learning • Knowledge of RLHF, DPO, model extraction, or related AI training and evaluation concepts • Cybersecurity experience involving penetration testing, exploit development, or reverse engineering • Experience testing conversational AI systems • Previous experience producing structured human data for AI evaluation
• Fully remote work • Flexible scheduling • Competitive hourly compensation of $45–$60 per hour • Weekly payments via Stripe or Wise • Participation in higher-sensitivity projects is optional • Projects may be extended, shortened, or adjusted depending on scope and performance
Apply Now🔥 1 hour ago
Energy sector expert evaluating AI-generated forecasts for Lightly AG, an ETH and HSG spin-off building machine-learning and computer-vision technology. Assessing forecast accuracy and response quality using a provided rubric.
🇺🇸 United States – Remote
⏳ Contract/Temporary
🟢 Junior
🟡 Mid-level
🤖 Artificial Intelligence
🚫👨🎓 No degree required
🔥 18 hours ago
AI Native Engineer building Claude Code agents, skills, and MCP integrations. Developing AWS-backed harness components for retrieval, tools, evaluation, and observability.
🇺🇸 United States – Remote
💵 $70 - $100 / hour
⏳ Contract/Temporary
🟡 Mid-level
🟠 Senior
🤖 Artificial Intelligence
🕒 Yesterday
STEM experts creating and solving PhD-level academic problems for Terac’s paid AI reasoning research studies. Reviewing AI-generated technical responses for logical and factual errors.
🇺🇸 United States – Remote
💵 $100 / hour
⏳ Contract/Temporary
🟡 Mid-level
🟠 Senior
🤖 Artificial Intelligence
🕒 Yesterday
STEM experts creating graduate-to-PhD-level problems and exact solutions for Terac’s AI training research. Reviewing AI responses and identifying scientific reasoning errors.
🇺🇸 United States – Remote
💵 $100 / hour
⏳ Contract/Temporary
🟡 Mid-level
🟠 Senior
🤖 Artificial Intelligence
🕒 Yesterday
Evaluating Hippocratic AI’s clinical agents for symptom triage, diabetes medication titration, and travel health guidance. Providing safety-focused feedback to improve AI-driven healthcare tools.
🇺🇸 United States – Remote
💰 $50M Seed Round on 2023-05
⏳ Contract/Temporary
🟡 Mid-level
🟠 Senior
🤖 Artificial Intelligence