
2 - 10 employees
š¤ B2B
š¼ Consulting
B2B ⢠Consulting
24-MAG is a commercial strategy and execution firm that helps B2B organizations design and implement systems, workflows, and operating rhythms for sales, client management, and cross-functional projects. They focus on transforming scattered processes into aligned, measurable, and scalable commercial functionsācovering pipeline structure, account management frameworks, and operational discipline for teams seeking efficient, intentional growth.
š„ 0 minutes ago
š½ New York ā Remote
šµ $45 - $60 / hour
ā³ Contract/Temporary
š” Mid-level
š Senior
š¤ Artificial Intelligence
š£ļøšøšŖ Swedish Required
Improve your chances of getting an interview by checking your resume score before you apply.

2 - 10 employees
š¤ B2B
š¼ Consulting
B2B ⢠Consulting
24-MAG is a commercial strategy and execution firm that helps B2B organizations design and implement systems, workflows, and operating rhythms for sales, client management, and cross-functional projects. They focus on transforming scattered processes into aligned, measurable, and scalable commercial functionsācovering pipeline structure, account management frameworks, and operational discipline for teams seeking efficient, intentional growth.
⢠Conduct adversarial testing of conversational AI models and agents ⢠Develop jailbreaks, prompt-injection scenarios, misuse cases, and multi-turn manipulation strategies ⢠Probe models for vulnerabilities missed by automated evaluation systems ⢠Test model behaviour across diverse conversational and adversarial scenarios ⢠Apply systematic testing frameworks and established evaluation methodologies ⢠Identify and classify model failures and safety vulnerabilities ⢠Evaluate bias, misinformation, misuse, and potentially harmful model behaviours ⢠Identify recurring or systemic patterns across model responses ⢠Assess severity, reproducibility, and practical significance of vulnerabilities ⢠Produce high-quality human evaluation data from red teaming activities ⢠Annotate model failures and categorise vulnerabilities ⢠Create structured attack cases and supporting evaluation materials ⢠Maintain consistency across repeated assessments and datasets ⢠Document adversarial scenarios clearly and reproducibly ⢠Produce reports, datasets, and structured findings for technical teams ⢠Explain identified risks to technical and non-technical stakeholders ⢠Record testing methodology, model behaviour, and relevant failure patterns ⢠Contribute to broader evaluation coverage across models and use cases
⢠Native-level fluency in both English and Swedish is required ⢠Prior experience with AI red teaming, adversarial AI evaluation, cybersecurity, or socio-technical system testing ⢠Strong understanding of conversational AI systems and model failure modes ⢠Experience developing structured adversarial tests and evaluation frameworks ⢠Ability to identify subtle vulnerabilities and recurring behavioural patterns ⢠Strong analytical reasoning and written communication skills ⢠Ability to document findings clearly and reproducibly ⢠Comfort working across changing projects, scenarios, and evaluation frameworks ⢠Background in computer science, cybersecurity, artificial intelligence, machine learning, linguistics, behavioural science, or a related discipline may be helpful ⢠Equivalent professional experience in AI safety, adversarial testing, security, or structured model evaluation may also be considered ⢠Practical red teaming experience is particularly valuable ⢠Familiarity with adversarial machine learning ⢠Knowledge of RLHF, DPO, model extraction, or related AI training and evaluation concepts ⢠Cybersecurity experience involving penetration testing, exploit development, or reverse engineering ⢠Experience testing conversational AI systems ⢠Native-level English and Swedish proficiency required for the engagement ⢠Work is as an independent contractor
⢠Fully remote work ⢠Flexible scheduling ⢠Competitive hourly compensation ⢠Weekly payments via Stripe or Wise ⢠Participation in higher-sensitivity projects is optional ⢠Flexible remote consulting work ⢠Projects may be extended, shortened, or adjusted depending on scope and performance
Apply Nowš„ 1 hour ago
Energy sector expert evaluating AI-generated forecasts for Lightly AG, an ETH and HSG spin-off building machine-learning and computer-vision technology. Assessing forecast accuracy and response quality using a provided rubric.
šŗšø United States ā Remote
ā³ Contract/Temporary
š¢ Junior
š” Mid-level
š¤ Artificial Intelligence
š«šØāš No degree required
š„ 18 hours ago
AI Native Engineer building Claude Code agents, skills, and MCP integrations. Developing AWS-backed harness components for retrieval, tools, evaluation, and observability.
šŗšø United States ā Remote
šµ $70 - $100 / hour
ā³ Contract/Temporary
š” Mid-level
š Senior
š¤ Artificial Intelligence
š Yesterday
STEM experts creating and solving PhD-level academic problems for Teracās paid AI reasoning research studies. Reviewing AI-generated technical responses for logical and factual errors.
šŗšø United States ā Remote
šµ $100 / hour
ā³ Contract/Temporary
š” Mid-level
š Senior
š¤ Artificial Intelligence
š Yesterday
STEM experts creating graduate-to-PhD-level problems and exact solutions for Teracās AI training research. Reviewing AI responses and identifying scientific reasoning errors.
šŗšø United States ā Remote
šµ $100 / hour
ā³ Contract/Temporary
š” Mid-level
š Senior
š¤ Artificial Intelligence
š Yesterday
Evaluating Hippocratic AIās clinical agents for symptom triage, diabetes medication titration, and travel health guidance. Providing safety-focused feedback to improve AI-driven healthcare tools.
šŗšø United States ā Remote
š° $50M Seed Round on 2023-05
ā³ Contract/Temporary
š” Mid-level
š Senior
š¤ Artificial Intelligence