
2 - 10 employees
π€ Artificial Intelligence
βοΈ SaaS
π’ Enterprise
Artificial Intelligence β’ SaaS β’ Enterprise
Polymath is a data lab focused on increasing the reliability and autonomy of AI agents by building simulation environments where agents can train and be evaluated. The company develops environments, applications, services, data, tasks, verifiers, and agents to enable long-horizon, low-supervision agent performance, and collaborates with leading model labs. Polymath's team comprises researchers, engineers, and operators working on safe, highly capable AI systems. It is backed by Base10, Y Combinator, and other investors.
π₯ 7 minutes ago
Improve your chances of getting an interview by checking your resume score before you apply.

2 - 10 employees
π€ Artificial Intelligence
βοΈ SaaS
π’ Enterprise
Artificial Intelligence β’ SaaS β’ Enterprise
Polymath is a data lab focused on increasing the reliability and autonomy of AI agents by building simulation environments where agents can train and be evaluated. The company develops environments, applications, services, data, tasks, verifiers, and agents to enable long-horizon, low-supervision agent performance, and collaborates with leading model labs. Polymath's team comprises researchers, engineers, and operators working on safe, highly capable AI systems. It is backed by Base10, Y Combinator, and other investors.
β’ Collaborate on a research project focused around frontier benchmarks and environments for long-horizon AI agents β’ Identify failure modes in frontier models β’ Develop rigorous benchmarks that evaluate how well frontier agents perform on complex, realistic tasks requiring long-horizon reasoning and tool use in dynamic environments β’ Train autonomous agents that can reason, plan, and act over extended time horizons
β’ Are currently pursuing an MS or PhD program in Computer Science or a related field β’ Have experience with reinforcement learning, benchmarking frontier models, or model post-training β’ Have experience with systems engineering and can write production-quality code β’ Have a strong track record of publications β’ Have high agency, move quickly, and enjoy working on open-ended research problems
β’ Health insurance β’ Flexible work arrangements β’ Professional development
Apply Nowπ 2 days ago
Experienced Designers reviewing design artifacts and drafting evaluation criteria for AI research. Contributing to the development of AI models that analyze visual design.
πΊπΈ United States β Remote
π΅ $30 / hour
β³ Contract/Temporary
π‘ Mid-level
π Senior
π§ AI Research Scientist