Search Remote Jobs

Evaluations Team Lead

Job not on LinkedIn

đŸ”„ 1 minute ago

đŸ‡ȘđŸ‡ș Europe – Remote

⏰ Full Time

🟠 Senior

đŸ‘» Ghost score 12%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Fundamental

Fundamental

51 - 200 employees

Founded 2024

đŸ€– Artificial Intelligence

🏱 Enterprise

☁ SaaS

Artificial Intelligence ‱ Enterprise ‱ SaaS

Fundamental is an enterprise AI company that builds large tabular models (LTMs) such as NEXUS, pre-trained on billions of tables to detect patterns and predict outcomes from structured data. The company offers an enterprise-grade predictive analytics platform that can be deployed with minimal code or integrated deeply with cloud partners like AWS, emphasizing privacy, security, and scalability. Born from academic research and backed by major investors, Fundamental targets large organizations seeking to extract foresight from their databases and deploy predictive models at cloud scale.

📋 Description

‱ Lead the team responsible for measuring NEXUS, Fundamental's Large Tabular Model ‱ Build and maintain a shared evaluation platform for engineering, research, and Applied AI ‱ Continuously add and maintain models and curated datasets for independent benchmarking and investigation ‱ Define standards for metrics, data splits, leakage prevention, calibration, uncertainty, and benchmark contamination ‱ Build reproducible pipelines with versioned inputs and artifacts ‱ Integrate regression checks into research and release workflows ‱ Benchmark NEXUS against competing approaches using fair tuning budgets, data access, compute, and latency protocols ‱ Support Applied AI customer proof-of-concept evaluations with tooling, methodological guidance, and analysis ‱ Measure predictive quality, latency, and cost across deployment configurations, task types, and dataset characteristics ‱ Convert findings into research priorities, release recommendations, and evidence-backed customer improvement plans ‱ Hire and develop a small team, set priorities, and remain hands-on with code and experimental design

🎯 Requirements

‱ Experience owning evaluation for tabular ML systems used in production or consequential customer decisions ‱ Strong statistical judgment: choosing metrics and validation schemes, estimating uncertainty, comparing models across datasets, and accounting for repeated experimentation ‱ Practical experience finding leakage in preprocessing, feature construction, joins, temporal dependencies, and related entities across splits ‱ Strong Python and SQL skills ‱ Familiarity with scikit-learn and gradient-boosted trees ‱ Experience building reliable ML tooling or platforms used by other teams ‱ Experience designing fair model comparisons, including hyperparameter search, resource budgets, and end-to-end latency measurement ‱ Prior people management experience, including hiring, technical coaching, and performance feedback, while remaining technically involved ‱ Clear written and spoken communication with researchers, engineers, and customer data scientists ‱ Experience evaluating tabular foundation models or AutoML systems (nice to have) ‱ Experience measuring how model optimisations affect predictive quality and inference performance (nice to have) ‱ Experience with relational or multi-table data, and Snowflake or Databricks environments (nice to have) ‱ Experience evaluating automated or agent-driven ML workflows, including failures that aggregate metrics can hide (nice to have)

đŸ–ïž Benefits

‱ Competitive compensation with salary and equity ‱ Comprehensive health coverage for you and your dependents ‱ Paid parental leave for all new parents, inclusive of adoptive and surrogate journeys ‱ Relocation support for employees moving to join the team in one of our office locations ‱ A mission-driven, low-ego culture that values diversity of thought, ownership, and bias toward action ‱ Fully remote role in Europe or Israel ‱ Monthly company onsites in Barcelona, with 4 days of travel

Apply Now

Similar Jobs

🕒 July 25

Adhoq

11 - 50

đŸ’Œ Consulting

📩 Logistics

📣 Marketing

Team Lead driving media buying strategies and mentoring a team in a global performance marketing agency. Focused on scaling profitable campaigns across multiple verticals and geos.

đŸ‡ȘđŸ‡ș Europe – Remote

⏰ Full Time

🟠 Senior

🕒 June 25

TSMG Holding

51 - 200

đŸ’Œ Consulting

📩 Logistics

đŸ€– Artificial Intelligence

Team Lead coordinating field data collection projects for TSMG across EU countries. Managing teams, schedules, and ensuring quality performance in urban imaging initiatives.

đŸ‡ȘđŸ‡ș Europe – Remote

⏰ Full Time

🟠 Senior