Software Engineer – Benchmarking

Job not on LinkedIn

🔥 0 minutes ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Office Hours

Office Hours

11 - 50 employees

Founded 2020

💼 Consulting

🏥 Healthcare

📣 Marketing

💰 $5M Seed Round - Office Hours on 2021-02

Consulting • Healthcare • Marketing

Office Hours is a platform that connects professionals across technology and life sciences with companies, researchers, and product teams to earn money by sharing expertise. Experts participate in paid 30–60 minute consultations, short surveys, and AI model-training tasks (data labeling, RLHF, expert evaluation), while the platform handles scheduling, payments, and opportunity matching. It operates as a global marketplace and SaaS-like service enabling domain experts to monetize knowledge and help shape AI and product decisions.

📋 Description

• Prepare and maintain benchmark datasets, including cleaning, preparation, conversion into runnable formats, validation, and ongoing maintenance • Build and maintain infrastructure and evaluation pipelines across model APIs and terminal agents • Create lightweight, containerized environments and viewers for tasking and tool-use evaluation • Support fine-tuning of small open-source LLMs and compare baseline with post-training performance • Develop scoreboards and leaderboards showing model-, benchmark-, task-, domain-, and rubric-level performance • Build analysis tools to identify failure modes, compare models and agent scaffolds, and track capability changes over time • Collaborate with researchers and engineers to ensure evaluation data and outputs are accurate, consistent, and integrated into published work • Build HTML viewers and internal tools for task authoring, review, quality control, and structured data collection

🎯 Requirements

• 4+ years of professional experience building and maintaining complex systems • Strong Python skills • Ability to write robust, maintainable code and work deeply in existing codebases and infrastructure • Experience preparing, cleaning, and maintaining datasets • Experience with Docker and reproducible execution environments • Ability to collaborate with researchers and scientists and translate methodology into working systems • Hands-on AI evaluation experience or experience with Harbor, Terminal-Bench, or Inspect is a strong plus • Experience with Python, model APIs, agent/evaluation frameworks, and custom evaluation tooling • Experience with terminal agents and open-source models via the Hugging Face ecosystem and PyTorch • Experience with Docker • Experience with React, Next.js, and Tailwind • Experience with GitHub, Slack, Notion, and Linear • Fine-tuning or post-training open-source LLM experience, or other hands-on machine learning work, is bonus experience • Experience with agentic, multi-turn, long-context, or tool-use evaluation is bonus experience • Experience validating LLM-as-judge or rubric-based grading setups is bonus experience • Background or strong interest in a scientific or technical domain is bonus experience • Experience building data-heavy dashboards, leaderboards, or visualizations is bonus experience • Open-source contributions or published work related to benchmarks and measurement is bonus experience

🏖️ Benefits

• Competitive salary and equity • Medical, dental, and vision coverage • 401(k) • Monthly wellness and fitness stipend • Paid time off policy, along with company holidays • Annual company off-sites (Tahoe, Mendocino, Mexico City, San Diego, Park City) • Parent-friendly policies • Remote flexibility • Paid family leave

Apply Now

Similar Jobs

🔥 12 minutes ago

Kong Inc.

201 - 500

💼 Consulting

📦 Logistics

🔌 API

Software Engineer building Golang IAM microservices for Kong’s cloud API platform. Developing secure, scalable authentication and authorization services for Kong Konnect and Developer Portal.

🔥 17 minutes ago

SecurityScorecard

501 - 1000

💼 Consulting

🏥 Healthcare

🛡️ Insurance

GTM Engineer building AI-powered Salesforce workflows and integrations for SecurityScorecard, a cybersecurity ratings company. Improving revenue-team automation, scoring, and data systems across the go-to-market organization.

🔥 17 minutes ago

SecurityScorecard

501 - 1000

💼 Consulting

🏥 Healthcare

🛡️ Insurance

GTM Engineer building AI-powered Salesforce workflows, integrations, and analytics for SecurityScorecard, a global cybersecurity ratings company. Supporting sales, marketing, and revenue operations with scalable automation and AI agents.

🔥 21 minutes ago

UpdatePromise

51 - 200

💼 Consulting

📦 Logistics

📣 Marketing

Senior Software Engineer taking ownership of backend architecture, operations, security, and automotive DMS/OEM integrations for PromisePay's software platform.

🔥 57 minutes ago

Attain Talent

1 - 10

🎯 Recruiter

🤝 Non-profit

📚 Education

Full Stack Engineer designing secure, scalable architectures for Attain Talent’s federal modernization clients. Leading cloud, API, integration, and enterprise platform solutions with engineering and DevSecOps teams.