Mercor Research Fellowship – APEX

Job not on LinkedIn

🔥 14 minutes ago

🏄 California – Remote

infoinfo

💵 $40k - $80k / year

⏳ Contract/Temporary

🟡 Mid-level

🟠 Senior

🧬 Research Scientist

👻 Ghost score 0%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Mercor

Mercor

51 - 200 employees

Founded 2023

💼 Consulting

📦 Logistics

📣 Marketing

Consulting • Logistics • Marketing

Mercor is a San Francisco-based online platform that connects companies with remote, paid AI and human-data workers for tasks like annotation, captioning, evaluation, and specialized consulting. The service operates as a marketplace offering role listings, daily payouts, and tools for managing data pipelines, incentives, and hiring workflows focused on AI productivity and research. Mercor targets businesses needing scalable human-in-the-loop data and AI-related talent, while providing candidates with gig-style opportunities and resources.

📋 Description

• Propose and scope a new benchmark or evaluation technique in an under-covered APEX domain or a meaningfully harder version of an existing one • Design task specifications and grading rubrics with Mercor’s vetted domain experts • Build and validate benchmarks through task pilots, scoring calibration, and stress-testing for contamination and gameable shortcuts • Run frontier models against benchmarks and analyze failure modes • Publish results as a paper, open dataset, APEX leaderboard, or methodology adopted internally • Partner with Mercor’s research and engineering teams to incorporate findings into APEX’s public benchmark family • Work directly with the APEX research team and access enterprise evaluation problems from Fortune 500 and frontier-lab partners

🎯 Requirements

• Genuine interest in evaluation as a research discipline • Background in CS, ML, statistics, or an adjacent field such as measurement, psychometrics, HCI, or social science • Specific, well-scoped idea for a benchmark or evaluation technique to build • Comfortable working in a startup environment with fast iteration and less hand-holding than an academic lab • Able to commit at least 20 hours/week for the duration of the fellowship • Minimum commitment of 30 hours/week; full-time preferred • Bonus: experience with agentic evaluation, RL environments, or domain expertise in law, finance, medicine, or a scientific field • Must submit a specific benchmark or evaluation-methodology proposal

🏖️ Benefits

• Unlimited API credits • Dedicated budget for GPU compute • Paid expert/human-data time • Weekly 1:1 mentorship with a member of the APEX research team • Regular access to the broader research organization • Access to frontier model APIs • Access to Mercor’s internal evaluation infrastructure • Access to real enterprise evaluation problems from Mercor’s customers, where appropriate • Optional desk in Mercor’s San Francisco office • Introductions to Mercor’s network of researchers across frontier labs and academia • Standout fellows considered for a full-time offer on the APEX research team at the end of the fellowship

Apply Now

Similar Jobs

🕒 2 days ago

Arizona

201 - 500

📣 Marketing

📱 Media

Senior Research Scientist developing AI/NLP systems for scientific feasibility assessment at the University of Arizona. Leading reproducible research pipelines, empirical studies, publications, and student collaboration.

🕒 August 13

Mercor

51 - 200

Biology Research Scientist validating UniProt protein targets for Mercor’s AI drug-discovery partner. Reviewing literature and assays to improve large-scale bioactivity databases.

🇺🇸 United States – Remote

💵 $50 - $70 / hour

🔥 Funding within the last year

💰 $350M Series C - Mercor on 2025-10

⏳ Contract/Temporary

🟢 Junior

🟡 Mid-level

🧬 Research Scientist

🕒 August 7

24-MAG

2 - 10

🤝 B2B

💼 Consulting

Machine learning research scientist advancing image and language models through remote consulting. Conducting experiments in robustness, efficiency, generative modelling, multilingual learning, and post-training.

🕒 July 29

Cystems Logic

51 - 200

Applied Scientist optimizing allocation models for retail inventory. Collaborating on mathematical optimization in a remote contract position with hands-on responsibilities.

🕒 July 28

Mento

51 - 200

☁️ SaaS

👥 HR Tech

🤝 B2B

Research Engineer/Research Scientist Coach providing research coaching to Mento members through virtual sessions. Seeking talent with AI/ML expertise and a strong academic background to guide members to fulfill their career goals.