
2 - 10 employees
🤝 B2B
💼 Consulting
B2B • Consulting
24-MAG is a commercial strategy and execution firm that helps B2B organizations design and implement systems, workflows, and operating rhythms for sales, client management, and cross-functional projects. They focus on transforming scattered processes into aligned, measurable, and scalable commercial functions—covering pipeline structure, account management frameworks, and operational discipline for teams seeking efficient, intentional growth.
🔥 18 hours ago
Improve your chances of getting an interview by checking your resume score before you apply.

2 - 10 employees
🤝 B2B
💼 Consulting
B2B • Consulting
24-MAG is a commercial strategy and execution firm that helps B2B organizations design and implement systems, workflows, and operating rhythms for sales, client management, and cross-functional projects. They focus on transforming scattered processes into aligned, measurable, and scalable commercial functions—covering pipeline structure, account management frameworks, and operational discipline for teams seeking efficient, intentional growth.
• Architect self-contained reinforcement-learning environments for complex real-world tasks • Design reward functions, verifiers, evaluation logic, and supporting environment components • Structure environments for reliable experimentation and measurable model improvement • Translate research objectives into rigorous RL workflows • Ensure environments are reproducible, testable, and suitable for iterative model development • Design and scale episode pipelines and multi-component training processes • Build reproducible experimentation workflows for reinforcement-learning research • Develop systems for running, tracking, and analysing large-scale training experiments • Improve reliability and efficiency across RL training infrastructure • Build automated synthetic data-generation systems • Develop AI-driven evaluation and quality-assurance systems for grading, validation, and feedback • Establish automated feedback loops to improve training-data and model quality • Design verification systems that distinguish strong model behaviour from superficially plausible outputs • Fine-tune and optimise open-source reinforcement-learning and machine-learning models • Develop benchmarking frameworks measuring capability, robustness, and data quality • Analyse model behaviour across internal and external evaluation environments • Contribute to the development, release, and interpretation of research evaluations and benchmark results • Operate across research experimentation and production-oriented technical implementation • Adapt research priorities, evaluation systems, and experimentation workflows as project requirements evolve
• Deep professional or research experience in reinforcement learning • Strong understanding of RL environment design, reward structures, training dynamics, and evaluation • Demonstrated experience building and scaling RL systems, training pipelines, or experimentation frameworks • Strong experience with automation and synthetic data-generation workflows • Familiarity with automated evaluation, model validation, and quality-assurance systems • Experience fine-tuning and evaluating open-source machine-learning models • Strong technical writing and communication skills • Ability to operate effectively in fast-paced, research-driven, and highly collaborative environments • Experience publishing benchmarks, evaluations, or research artifacts is advantageous • Familiarity with modern evaluation ecosystems and benchmarking frameworks is beneficial • Experience with scalable infrastructure supporting large-scale RL experimentation is strongly valued • Must work without using confidential or proprietary information belonging to any employer, client, institution, or other third party
• Fully remote work • Full-time engagement • Compensation of $400,000–$800,000/year • Remote consulting opportunities across technical, evaluation, and project-based workstreams
Apply Now🕒 September 2
Staff parser/compiler research engineer building high-performance constrained decoding systems. Enabling .txt’s reliable and safe language-model agents through grammars, parsers, and compilers.
🕒 July 28
Staff+ Forward Deployed Research Engineer at Sequen developing production-grade systems for AI-driven ranking and recommendations. Working at the intersection of AI research and customized deployment inside enterprise environments.
🕒 July 28
Research Engineer working on AI world modeling and multimodal foundational models at innovative tech company. Collaborating across research efforts and bringing models to production.
🕒 November 8, 2025
Staff Research Engineer developing techniques for improving speed and efficiency of AI models at Cohere. Join a diverse team focused on advanced AI research and model optimization.