Research Scientist, Benchmarks & Evals

Job not on LinkedIn

🔥 0 minutes ago

🇺🇸 United States – Remote

💵 $150k - $250k / year

⏰ Full Time

🟡 Mid-level

🟠 Senior

🧬 Research Scientist

👻 Ghost score 0%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Kled AI

Kled AI

11 - 50 employees

Founded 2025

🤖 Artificial Intelligence

🏪 Marketplace

🤝 B2B

Artificial Intelligence • Marketplace • B2B

Kled AI is a research-backed data marketplace and mobile platform that sources, curates, and licenses opt-in human datasets to train and improve AI models. Through a contributor app that pays users for photos, videos, and other human-generated data, Kled aggregates high-utility, verifiable datasets for AI companies, governments, and research institutions, and publishes research demonstrating dataset improvements for model fidelity. The company focuses on solving AI data scarcity by combining marketplace services, contributor tools, automated quality checks, and enterprise-facing dataset products.

📋 Description

• Design and ship public benchmarks for image, video and document models • Build private test sets from real, opt-in data that no model has trained on • Run data-value experiments by training models with and without Kled data and measuring changes • Design human evaluation studies and automatic metrics • Publish leaderboards and reports trusted by researchers at frontier labs • Identify model shortcomings and advise what data to collect next • Work directly with founders and lab buyers to turn results into data deals

🎯 Requirements

• 3+ years ML research or research engineering experience • A benchmark or evaluation you built that other people actually used • Strong hands-on experience training and fine-tuning image or video models (PyTorch) • Deep understanding of experimental design and statistics (controls, ablations, human studies) • Experience running experiments end to end, from GPUs to final report • Clear writing for both researchers and buyers • Publications at NeurIPS, ICML, ICLR or CVPR (datasets and benchmarks especially) (bonus) • Experience on an evaluation or data team at an AI lab (bonus) • Experience selling or licensing data to AI labs (bonus) • Experience with data valuation or data attribution research (bonus) • Python / PyTorch • Open-weight image and video models • Cloud GPUs • PostgreSQL (Supabase) • S3 storage

🏖️ Benefits

• $350,000 – $750,000 equity • 0.25% – 1% equity

Apply Now

Similar Jobs

🕒 2 days ago

DECA

11 - 50

💼 Consulting

💸 Finance

🤝 B2B

Senior Principal Scientist shaping US regulatory strategy for Dechra’s veterinary drugs and biologics. Guiding FDA/USDA submissions, product approvals, and regulatory innovation.

🕒 3 days ago

Zillow

5001 - 10000

🏠 Real Estate

🛍️ eCommerce

👥 B2C

Applied Scientist developing econometric and machine-learning housing forecasts for Zillow. Building regional models, production forecasts, and market stress-testing scenarios.

🕒 5 days ago

Index Analytics LLC

51 - 200

💼 Consulting

🏥 Healthcare

📦 Logistics

Senior researcher analyzing Medicaid and CHIP data for Index Analytics, a federal healthcare consulting firm. Building predictive models, analytic products, and policy insights.

🕒 5 days ago

Lucas James Talent Partners

51 - 200

💼 Consulting

📣 Marketing

🎯 Recruiter

Research Scientist advancing AI safety systems in NLP, computer vision, signal processing, and AI reasoning. Developing rigorous evaluations and digital safety tools for nonprofit UL Research Institutes.

🕒 6 days ago

MSD

10,000+ employees

🏥 Healthcare

📦 Logistics

🧬 Biotechnology

Senior Principal Scientist leading global regulatory strategies for Merck’s immunology drug and biologic programs. Guiding submissions, health-authority interactions, labeling, and product approvals worldwide.