Search Remote Jobs

Member of Technical Staff, Coding Research

Job not on LinkedIn

🔥 17 hours ago

🗽 New York – Remote

infoinfo

💵 $600k - $1.3M / year

⏰ Full Time

🔴 Lead

🖥 Software Engineer

👻 Ghost score 0%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of 24-MAG

24-MAG

2 - 10 employees

🤝 B2B

💼 Consulting

B2B • Consulting

24-MAG is a commercial strategy and execution firm that helps B2B organizations design and implement systems, workflows, and operating rhythms for sales, client management, and cross-functional projects. They focus on transforming scattered processes into aligned, measurable, and scalable commercial functions—covering pipeline structure, account management frameworks, and operational discipline for teams seeking efficient, intentional growth.

📋 Description

• Design and own evaluation frameworks for advanced coding agents • Develop benchmark specifications, scoring methodologies, rubrics, and quality standards • Establish rigorous methods for measuring coding-model performance across diverse software-engineering tasks • Define objective criteria for correctness, reasoning quality, robustness, and task completion • Maintain methodological rigour, reproducibility, and consistency across evaluation workflows • Develop high-quality datasets, golden examples, and structured evaluation protocols • Design technical tasks for reliable assessment of frontier coding systems • Build data and evaluation workflows supporting model development and iterative improvement • Identify benchmark or dataset coverage gaps and develop new evaluation categories • Analyse coding-agent behaviour and identify systematic weaknesses, failure modes, and performance limitations • Investigate incorrect reasoning, implementation errors, tool-use failures, and incomplete task execution • Translate findings into recommendations for model training and evaluation • Design experiments testing hypotheses about coding-model capabilities • Build tooling and infrastructure for large-scale experimentation, data generation, review workflows, and evaluation pipelines • Automate technical processes to improve evaluation efficiency and research velocity • Collaborate with researchers, engineers, and applied AI teams • Contribute to technical reports, benchmark studies, research documentation, and external-facing research initiatives • Communicate complex technical findings to specialist and broader technical audiences

🎯 Requirements

• Strong software-engineering background with expertise in Python, C++, or comparable programming languages • Minimum of 3 years of experience in software engineering, machine learning, AI research, evaluation, or a related technical discipline • Experience designing, reviewing, or validating technical assessments, benchmarks, coding tasks, or evaluation methodologies • Familiarity with large language models, coding agents, reinforcement learning, model evaluation, or related AI systems • Proven ability to build tooling, automate workflows, and improve technical processes through systematic experimentation • Strong analytical skills and ability to investigate complex model behaviour and technical failure modes • Excellent written and verbal communication skills • Ability to operate effectively in fast-moving research environments with significant ambiguity and evolving priorities • Experience with frontier AI systems, coding agents, or model-evaluation research is advantageous • Experience designing benchmarks or datasets for machine-learning systems at scale is strongly valued • Familiarity with agentic workflows, tool use, reinforcement learning, or post-training methodologies is beneficial • Publications, open-source contributions, or demonstrated technical leadership in AI, machine learning, or software engineering are advantageous • Work must be completed without using confidential or proprietary information belonging to any employer, client, institution, or other third party

🏖️ Benefits

• Fully remote work • Full-time engagement • Opportunity to contribute to frontier AI research and development • Collaboration with researchers, engineers, and applied AI teams • Opportunity to contribute to technical reports, benchmark studies, research documentation, and external-facing research initiatives

Apply Now

Similar Jobs

🔥 17 hours ago

Conduent

10,000+ employees

🤝 B2B

☁️ SaaS

🏢 Enterprise

Informatica Developer building secure healthcare EDI solutions with PowerCenter and B2B Data Transformation. Supporting ANSI X12 5010 platforms for Conduent’s enterprise and government clients.

🇺🇸 United States – Remote

💵 $80.1k - $104k / year

💰 $93M Post-IPO Debt - Conduent on 2025-08

⏰ Full Time

🟠 Senior

🔴 Lead

🖥 Software Engineer

🕒 Yesterday

YA Group

501 - 1000

💼 Consulting

🏗️ Construction

🛡️ Insurance

Director leading forensic engineering investigations for YA Group’s building consulting practice. Mentoring engineers, inspecting damaged structures, analyzing failures, and delivering defensible technical reports.

🕒 Yesterday

Sparibis

11 - 50

💼 Consulting

🔒 Cybersecurity

🏢 Enterprise

Software Engineer developing Java web applications and Azure Cloud systems for Sparibis, a professional solutions firm. Building CI/CD pipelines, integrating software, and maintaining secure production environments.

🕒 Yesterday

ShorePoint Inc

1 - 10

💼 Consulting

🏥 Healthcare

📦 Logistics

Archer GRC Developer configuring RSA Archer for federal cybersecurity, risk, and compliance initiatives at ShorePoint. Building workflows, integrations, dashboards, and automation aligned with NIST and FISMA.

🕒 Yesterday

Conduent

10,000+ employees

🏥 Healthcare

📦 Logistics

💼 Consulting

Informatica Developer building secure healthcare EDI solutions for Conduent’s mission-critical client services. Designing PowerCenter mappings, workflows, and B2B Data Transformation applications.