
2 - 10 employees
đ¤ B2B
đź Consulting
B2B ⢠Consulting
24-MAG is a commercial strategy and execution firm that helps B2B organizations design and implement systems, workflows, and operating rhythms for sales, client management, and cross-functional projects. They focus on transforming scattered processes into aligned, measurable, and scalable commercial functionsâcovering pipeline structure, account management frameworks, and operational discipline for teams seeking efficient, intentional growth.
đĽ 17 hours ago
đ˝ New York â Remote
đľ $600k - $1.3M / year
â° Full Time
đ´ Lead
đĽ Software Engineer
đť Ghost score 0%
Improve your chances of getting an interview by checking your resume score before you apply.

2 - 10 employees
đ¤ B2B
đź Consulting
B2B ⢠Consulting
24-MAG is a commercial strategy and execution firm that helps B2B organizations design and implement systems, workflows, and operating rhythms for sales, client management, and cross-functional projects. They focus on transforming scattered processes into aligned, measurable, and scalable commercial functionsâcovering pipeline structure, account management frameworks, and operational discipline for teams seeking efficient, intentional growth.
⢠Design and own evaluation frameworks for advanced coding agents ⢠Develop benchmark specifications, scoring methodologies, rubrics, and quality standards ⢠Establish rigorous methods for measuring coding-model performance across diverse software-engineering tasks ⢠Define objective criteria for correctness, reasoning quality, robustness, and task completion ⢠Maintain methodological rigour, reproducibility, and consistency across evaluation workflows ⢠Develop high-quality datasets, golden examples, and structured evaluation protocols ⢠Design technical tasks for reliable assessment of frontier coding systems ⢠Build data and evaluation workflows supporting model development and iterative improvement ⢠Identify benchmark or dataset coverage gaps and develop new evaluation categories ⢠Analyse coding-agent behaviour and identify systematic weaknesses, failure modes, and performance limitations ⢠Investigate incorrect reasoning, implementation errors, tool-use failures, and incomplete task execution ⢠Translate findings into recommendations for model training and evaluation ⢠Design experiments testing hypotheses about coding-model capabilities ⢠Build tooling and infrastructure for large-scale experimentation, data generation, review workflows, and evaluation pipelines ⢠Automate technical processes to improve evaluation efficiency and research velocity ⢠Collaborate with researchers, engineers, and applied AI teams ⢠Contribute to technical reports, benchmark studies, research documentation, and external-facing research initiatives ⢠Communicate complex technical findings to specialist and broader technical audiences
⢠Strong software-engineering background with expertise in Python, C++, or comparable programming languages ⢠Minimum of 3 years of experience in software engineering, machine learning, AI research, evaluation, or a related technical discipline ⢠Experience designing, reviewing, or validating technical assessments, benchmarks, coding tasks, or evaluation methodologies ⢠Familiarity with large language models, coding agents, reinforcement learning, model evaluation, or related AI systems ⢠Proven ability to build tooling, automate workflows, and improve technical processes through systematic experimentation ⢠Strong analytical skills and ability to investigate complex model behaviour and technical failure modes ⢠Excellent written and verbal communication skills ⢠Ability to operate effectively in fast-moving research environments with significant ambiguity and evolving priorities ⢠Experience with frontier AI systems, coding agents, or model-evaluation research is advantageous ⢠Experience designing benchmarks or datasets for machine-learning systems at scale is strongly valued ⢠Familiarity with agentic workflows, tool use, reinforcement learning, or post-training methodologies is beneficial ⢠Publications, open-source contributions, or demonstrated technical leadership in AI, machine learning, or software engineering are advantageous ⢠Work must be completed without using confidential or proprietary information belonging to any employer, client, institution, or other third party
⢠Fully remote work ⢠Full-time engagement ⢠Opportunity to contribute to frontier AI research and development ⢠Collaboration with researchers, engineers, and applied AI teams ⢠Opportunity to contribute to technical reports, benchmark studies, research documentation, and external-facing research initiatives
Apply NowđĽ 17 hours ago
Informatica Developer building secure healthcare EDI solutions with PowerCenter and B2B Data Transformation. Supporting ANSI X12 5010 platforms for Conduentâs enterprise and government clients.
đşđ¸ United States â Remote
đľ $80.1k - $104k / year
đ° $93M Post-IPO Debt - Conduent on 2025-08
â° Full Time
đ Senior
đ´ Lead
đĽ Software Engineer
đ Yesterday
Director leading forensic engineering investigations for YA Groupâs building consulting practice. Mentoring engineers, inspecting damaged structures, analyzing failures, and delivering defensible technical reports.
đşđ¸ United States â Remote
đ° Private equity on 2021-09
â° Full Time
đ´ Lead
đĽ Software Engineer
đ Yesterday
Software Engineer developing Java web applications and Azure Cloud systems for Sparibis, a professional solutions firm. Building CI/CD pipelines, integrating software, and maintaining secure production environments.
đ Yesterday
Archer GRC Developer configuring RSA Archer for federal cybersecurity, risk, and compliance initiatives at ShorePoint. Building workflows, integrations, dashboards, and automation aligned with NIST and FISMA.
đ Yesterday
Informatica Developer building secure healthcare EDI solutions for Conduentâs mission-critical client services. Designing PowerCenter mappings, workflows, and B2B Data Transformation applications.
đşđ¸ United States â Remote
đľ $80.1k - $104k / year
đ° Venture Round on 2009-01
â° Full Time
đ Senior
đ´ Lead
đĽ Software Engineer
đŚ H1B Visa Sponsor