
2 - 10 employés
🤝 B2B
💼 Conseil
B2B • Consulting
24-MAG est une entreprise spécialisée dans la stratégie commerciale et l'exécution qui aide les organisations B2B à concevoir et mettre en œuvre des systèmes, des flux de travail et des rythmes opérationnels pour les ventes, la gestion client et les projets transversaux. Elle se concentre sur la transformation de processus dispersés en fonctions commerciales alignées, mesurables et évolutives, couvrant la structure du pipeline, les cadres de gestion de comptes, et la discipline opérationnelle pour les équipes recherchant une croissance efficiente et intentionnelle.
🕒 il y a 25 jours
🗣️🇺🇸🇬🇧 Anglais requis
Améliorez vos chances d'obtenir un entretien en vérifiant votre score de CV avant de postuler.

2 - 10 employés
🤝 B2B
💼 Conseil
B2B • Consulting
24-MAG est une entreprise spécialisée dans la stratégie commerciale et l'exécution qui aide les organisations B2B à concevoir et mettre en œuvre des systèmes, des flux de travail et des rythmes opérationnels pour les ventes, la gestion client et les projets transversaux. Elle se concentre sur la transformation de processus dispersés en fonctions commerciales alignées, mesurables et évolutives, couvrant la structure du pipeline, les cadres de gestion de comptes, et la discipline opérationnelle pour les équipes recherchant une croissance efficiente et intentionnelle.
• Adversarial Model Evaluation • Probe frontier AI models across coding, machine learning, analysis, and multi-step agentic tasks • Identify subtle errors, vulnerabilities, edge cases, and misleadingly plausible outputs • Investigate situations where models appear capable while reaching incorrect or unsupported conclusions • Design reproducible experiments to isolate and validate model failure modes • Benchmark & Challenge Design • Convert observed model weaknesses into rigorous benchmark tasks • Develop challenges that are technically demanding while remaining fair and objectively assessable • Define clear task requirements, expected outcomes, and evaluation criteria • Ensure tasks require genuine reasoning rather than allowing shortcuts or superficial pattern matching • Failure Analysis & Documentation • Document findings with clear evidence, methodology, and reproducible steps • Explain why a model failed and which capabilities or assumptions contributed to the error • Produce detailed technical write-ups for researchers and task authors • Track recurring failure patterns across models, prompts, and evaluation environments • Task Strengthening & Research Collaboration • Work with task authors to close loopholes, grading gaps, and unintended shortcuts • Review benchmark tasks for ambiguity, exploitability, and evaluation reliability • Share insights with researchers and other specialists to improve benchmark coverage • Participate in iterative calibration, peer review, and task-refinement workflows
• At least 1 year of experience in research, research engineering, security, AI evaluation, or a related technical role • Demonstrated experience identifying vulnerabilities, edge cases, or failure modes in LLMs or ML systems • Background in red teaming, adversarial testing, security research, benchmark development, or rigorous model evaluation • Working proficiency in Python and Git • Ability to develop scripts, probes, and analyses independently • Strong familiarity with LLM capabilities, limitations, and evaluation techniques • Excellent written communication and technical documentation skills • Creativity, precision, and persistence when working through ambiguous research problems • Reliable availability for approximately 35 hours per week • A master's degree or PhD in a STEM field is highly relevant
Postuler Maintenant🕒 il y a 25 jours
Outreach Specialist connecting K-6 schools to the Book Blast program through presentations and meetings. Focused on scheduling with decision-makers to promote student home libraries.
🗣️🇺🇸🇬🇧 Anglais requis
🕒 il y a 28 jours
DERMS Solution Specialist deploying GridOS software for utility customers. Designing distributed-energy planning solutions, testing workflows, training users, and supporting delivery and sales.
🗣️🇺🇸🇬🇧 Anglais requis
🕒 il y a 28 jours
DERMS solution specialist deploying GridOS for utility distributed-energy planning. Advising utilities, designing customer solutions, and training users for cleaner grid operations.
🗣️🇺🇸🇬🇧 Anglais requis
🕒 il y a 28 jours
Visual Merchandising Specialist focused on store layout and product arrangement at Nouria Energy. Collaborating with merchandising and marketing teams to maximize sales and customer experience.
🗣️🇺🇸🇬🇧 Anglais requis
🕒 il y a 28 jours
Placement Coordinator responsible for transcript evaluations and assisting students and families with placement and graduation planning. Work remotely while collaborating with a team across various locations.
🗣️🇺🇸🇬🇧 Anglais requis