Evaluation Engineer

🕒 January 22

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Elicit

Elicit

11 - 50 employees

Founded 2023

📚 Education

🤖 Artificial Intelligence

⚕️ Healthcare Insurance

Education • Artificial Intelligence • Healthcare Insurance

Elicit is a tool designed to enhance academic research productivity by automating the analysis of research papers. It allows users to summarize papers, extract data, and synthesize findings quickly and efficiently, making the literature review process significantly faster and more organized. With access to over 125 million research papers, Elicit empowers researchers to discover relevant literature using natural language queries and offers features that streamline systematic reviews and meta-analyses.

📋 Description

• You'll build a comprehensive system that runs fast, is easy to use, and supports quickly building new evals: • Speed: You’ll build a lightning-fast basic evals infrastructure that schedules tasks to introduce practically no latency; and then you’ll figure out clever ways to solve the fundamental sources of latency (building a version of Elicit, running it on a query, and evaluating it using LMs) • Interfaces: ML engineers need evals to kick off automatically on relevant commits, with results they can see at a glance and drill into. Product managers need dashboards showing performance over time and what's going wrong in production. • Architecture: Your code must be well-architected so other team members and ML engineers can understand and build on it. An engineer starting on a new feature should be able to quickly add examples and run an eval. • We need to evaluate how well Elicit actually helps with decision-making in pharma, not just measure what's easy to measure. This requires encoding real knowledge about how pharma customers make decisions (for example, choosing appropriate gold standards). • You'll provide appropriate statistical tests and confidence intervals so we can trust our results. • In a typical month, expect to spend: • 60% working on the core eval platform • 15% working closely with the evals team to build and improve specific evals (e.g., an eval of our paper search within our systematic review flow) • 10% mentoring our evals engineering intern • The rest on learning how people interact with the eval system so you can make it work better for them, and understanding what our users want from Elicit so evals measure what matters

🎯 Requirements

• At least 3 years of experience as a professional software engineer, with demonstrated experience building complex backend systems (e.g., backend for a complex website, data pipelines, etc.) • Aptitude and interest in evaluating how Elicit helps with pharma decision-making. There's no particular experience you must have, but we'll evaluate your aptitude. • Knowledge of statistics (for e.g. calculating power and credence intervals for evals) • Experience with advanced Python (asyncio/trio and parallel processing strategies) • Front-end experience and strong UX sensibility (you'll be building dashboards). TypeScript experience is a plus. • Experience building developer tools (ML engineers are one of your most important clients) • Previous experience as a data engineer or working on AI infrastructure • Knowledge of pharma/biomed • Experience evaluating ML systems • Experience building language-model-based systems (helps with understanding Elicit and how to evaluate it)

🏖️ Benefits

• Flexible work environment: work from our office in Oakland or remotely with time zone overlap (between GMT and GMT-8), as long as you can travel for in-person retreats and coworking events • Fully covered health, dental, vision, and life insurance for you, generous coverage for the rest of your family • Flexible vacation policy, with a minimum recommendation of 20 days/year + company holidays • 401K with a 6% employer match • A new Mac + $1,000 budget to set up your workstation or home office in your first year, then $500 every year thereafter • $1,000 quarterly AI Experimentation & Learning budget, so you can freely experiment with new AI tools, take courses, purchase educational resources, or attend AI-focused conferences and events • A team administrative assistant who can help you with personal and work tasks

Apply Now

Similar Jobs

🕒 January 22

RESPEC

201 - 500

Lead Human Factors Engineer ensuring user-centric design for OCI's digital solutions. Focusing on reducing clinician burnout and enhancing system usability in healthcare settings.

🕒 January 21

Glean

11 - 50

🤖 Artificial Intelligence

🏢 Enterprise

⚡ Productivity

Founding Forward Deployed Engineer shaping transformational AI solutions for Glean's customers. Collaborating across teams to launch production platforms that deliver business value.

AWS

Azure

Cloud

Google Cloud Platform

Go

🕒 January 21

Terabase Energy

51 - 200

⚡ Energy

☁️ SaaS

Controls Engineer leading modeling, simulation, and testing for renewable energy systems. Overseeing project documentation, schedules, and continuous improvement in control processes.

🕒 January 16

Golden Prospects by YMP

11 - 50

🎯 Recruiter

👥 HR Tech

☁️ SaaS

Senior Atlassian Engineer managing Jira and Confluence tools for a leading AI data organization. Collaborating with executives to drive enterprise-wide transformations and best practices.

Ansible

Groovy

Python

🕒 January 15

Skyways

1 - 10

🚀 Aerospace

🚗 Transport

Designing and implementing flight control systems for autonomous aircraft at Skyways. Collaborating with various teams on real-world aviation innovations.

Python