
2 - 10 employees
š¤ B2B
š¼ Consulting
B2B ⢠Consulting
24-MAG is a commercial strategy and execution firm that helps B2B organizations design and implement systems, workflows, and operating rhythms for sales, client management, and cross-functional projects. They focus on transforming scattered processes into aligned, measurable, and scalable commercial functionsācovering pipeline structure, account management frameworks, and operational discipline for teams seeking efficient, intentional growth.
š„ 1 minute ago
š½ New York ā Remote
šµ $45 - $60 / hour
ā³ Contract/Temporary
š” Mid-level
š Senior
š¤ Artificial Intelligence
š£ļøš³š“ Norwegian Required
Improve your chances of getting an interview by checking your resume score before you apply.

2 - 10 employees
š¤ B2B
š¼ Consulting
B2B ⢠Consulting
24-MAG is a commercial strategy and execution firm that helps B2B organizations design and implement systems, workflows, and operating rhythms for sales, client management, and cross-functional projects. They focus on transforming scattered processes into aligned, measurable, and scalable commercial functionsācovering pipeline structure, account management frameworks, and operational discipline for teams seeking efficient, intentional growth.
⢠Conduct adversarial testing of conversational AI models and agents ⢠Develop jailbreaks, prompt-injection scenarios, misuse cases, and multi-turn manipulation strategies ⢠Probe models for vulnerabilities across diverse conversational and adversarial scenarios ⢠Identify and classify model failures and safety vulnerabilities ⢠Evaluate bias, misinformation, misuse, and potentially harmful model behaviours ⢠Assess vulnerability severity, reproducibility, and practical significance ⢠Produce and annotate human evaluation data from red teaming activities ⢠Create structured attack cases and evaluation materials ⢠Document adversarial scenarios, testing methodology, model behaviour, and failure patterns ⢠Produce reports, datasets, and structured findings for technical teams ⢠Explain identified risks to technical and non-technical stakeholders ⢠Contribute to broader evaluation coverage across models and use cases
⢠Native-level fluency in both English and Norwegian ⢠Prior experience with AI red teaming, adversarial AI evaluation, cybersecurity, or socio-technical system testing ⢠Strong understanding of conversational AI systems and model failure modes ⢠Experience developing structured adversarial tests ⢠Ability to identify subtle vulnerabilities and recurring behavioural patterns ⢠Strong analytical reasoning and written communication skills ⢠Ability to document findings clearly and reproducibly ⢠Comfort working across changing projects, scenarios, and evaluation frameworks ⢠Equivalent professional experience in AI safety, adversarial testing, security, or structured model evaluation may be considered ⢠Practical red teaming experience is particularly valuable ⢠Familiarity with adversarial machine learning ⢠Knowledge of RLHF, DPO, model extraction, or related AI training and evaluation concepts ⢠Cybersecurity experience involving penetration testing, exploit development, or reverse engineering ⢠Experience analysing abuse, harassment, misinformation, or other socio-technical risks ⢠Experience testing conversational AI systems ⢠Previous experience producing structured human data for AI evaluation
⢠Fully remote work ⢠Flexible scheduling ⢠Competitive hourly compensation ⢠Weekly payments via Stripe or Wise ⢠Participation in higher-sensitivity projects is optional ⢠Flexible remote consulting work ⢠Projects may be extended, shortened, or adjusted depending on scope and performance
Apply Nowš„ 1 hour ago
Energy sector expert evaluating AI-generated forecasts for Lightly AG, an ETH and HSG spin-off building machine-learning and computer-vision technology. Assessing forecast accuracy and response quality using a provided rubric.
šŗšø United States ā Remote
ā³ Contract/Temporary
š¢ Junior
š” Mid-level
š¤ Artificial Intelligence
š«šØāš No degree required
š„ 18 hours ago
AI Native Engineer building Claude Code agents, skills, and MCP integrations. Developing AWS-backed harness components for retrieval, tools, evaluation, and observability.
šŗšø United States ā Remote
šµ $70 - $100 / hour
ā³ Contract/Temporary
š” Mid-level
š Senior
š¤ Artificial Intelligence
š Yesterday
STEM experts creating and solving PhD-level academic problems for Teracās paid AI reasoning research studies. Reviewing AI-generated technical responses for logical and factual errors.
šŗšø United States ā Remote
šµ $100 / hour
ā³ Contract/Temporary
š” Mid-level
š Senior
š¤ Artificial Intelligence
š Yesterday
STEM experts creating graduate-to-PhD-level problems and exact solutions for Teracās AI training research. Reviewing AI responses and identifying scientific reasoning errors.
šŗšø United States ā Remote
šµ $100 / hour
ā³ Contract/Temporary
š” Mid-level
š Senior
š¤ Artificial Intelligence
š Yesterday
Evaluating Hippocratic AIās clinical agents for symptom triage, diabetes medication titration, and travel health guidance. Providing safety-focused feedback to improve AI-driven healthcare tools.
šŗšø United States ā Remote
š° $50M Seed Round on 2023-05
ā³ Contract/Temporary
š” Mid-level
š Senior
š¤ Artificial Intelligence