Senior Data Scientist – AI Evaluation & Quality

🔥 1 hour ago

🇱🇹 Lithuania – Remote

⏰ Full Time

🟠 Senior

📊 Data Scientist

👻 Ghost score 16%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Finom

Finom

501 - 1000 employees

Founded 2019

💳 Fintech

💸 Finance

🤝 B2B

Fintech • Finance • B2B

Finom is a fintech company offering comprehensive financial services for SMEs, freelancers, and companies. The platform provides tools for invoicing, expense management, and easy online account opening with added benefits like cashback and international IBAN accounts. Finom ensures data security through partnerships with banks like SolarisBank and Treezor, making it a fast and reliable service for business financial management.

📋 Description

• Join the AI Team driving Finom’s AI products and technology • Own and extend the offline evaluation suite across products, including datasets, judges, and metrics • Build and maintain online quality dashboards covering resolution rate, CSAT, thumbs up/down, LLM-as-judge signals, error rate, and latency • Close the production feedback loop by mining failure patterns from real traffic and turning them into regression cases • Propose fixes to Product and domain experts • Harden evaluation methodology, including judge stability and non-determinism handling • Translate metrics into decisions through weekly syncs and clear trade-offs • Collaborate closely with AI engineers, Product, and domain experts across the company • Work with Databricks, DeepEval, and Claude Code

🎯 Requirements

• Python and SQL; ability to build an analysis end-to-end • Solid foundation in statistics, including sampling, hypothesis testing, and variance • Analytical mindset focused on business questions • 3+ years in analyst/data scientist roles, including at least 1 year in a product context • Experience in quality analytics for ML systems (nice-to-have) • Hands-on experience evaluating LLM applications such as RAG, agents, tool use, and judges (nice-to-have) • Experience building LLM agents (nice-to-have) • AI-assisted coding as the default authoring environment • Familiarity or willingness to become fluent quickly with Claude Code for SQL, Python, analyses, dashboards, and internal scripts

🏖️ Benefits

• Continuous personal and professional development resources and opportunities • Flexibility to travel and work remotely or in a hybrid model across Europe • Stock Options Program available to every team member • Constant support and care • Modern, friendly, and eco-conscious corporate culture • Work & Swim Program: one month in a corporate apartment in Cyprus

Apply Now