Agent Quality, Evals Engineer

🕒 June 23

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Softgic

Softgic

51 - 200 employees

Founded 2011

💼 Consulting

🤝 B2B

🤖 Artificial Intelligence

Consulting • B2B • Artificial Intelligence

Softgic is a digital and cognitive transformation services company that provides custom software development, automation, cloud and cybersecurity, and emerging-technology solutions to business clients. With more than a decade of experience, Softgic offers specialized “labs” for Solutions Design (UX, Agile, process design), Build (coding, low-code, mobile, DevOps), Metaverse (VR/AR, blockchain, NFTs, virtual economy), Data (Big Data, BI, AI/ML), Automation (RPA, BPM) and Go Live (cloud, IT operations, cybersecurity). The firm emphasizes tailored, enterprise-focused services, dedicated teams, and project-based engagements to help clients modernize infrastructure, streamline processes, and adopt advanced technologies.

📋 Description

• Own the eval harness and quality gate from the beginning • Build and maintain the MVP eval harness, including golden tasks, exception tasks, scorecard metrics, and regression packs • Wire evaluations into CI so quality regressions fail builds and releases • Define and maintain release-gate thresholds with Product and the Tech Lead • Prepare for later adversarial and drift-testing expansion without overbuilding MVP scope • Use AI to generate candidate evaluation cases and failure hypotheses while validating generated tests • Establish the first reference agent's published scorecard and gated evaluation path • Automate golden and exception tests • Define measurable 'good enough to ship' criteria

🎯 Requirements

• Experience evaluating ML, LLM, or non-deterministic systems • Strong test and benchmark design capability • Comfort working with noisy metrics, thresholds, and probabilistic behavior • Good scripting and automation skills • Uses AI to generate candidate eval cases and failure hypotheses, while distinguishing generated tests from validated quality • Treats AI quality as an operating system rather than a QA afterthought

Apply Now

Similar Jobs

🕒 June 22

Lumos

51 - 200

🤝 Non-profit

🤲 Charity

🌍 Social Impact

OSP Project Engineer at Lumos responsible for planning and preparing construction drawings for fiber internet infrastructure. Designing for optimal use of communications facilities in a rapidly growing company.

🕒 June 22

Ensono

1001 - 5000

💼 Consulting

Senior IAM Engineer overseeing the operational maintenance and expansion of ForgeRock IAM platform. Ensuring high availability and optimal performance while developing custom scripts and configurations.

🕒 June 22

Terabase Energy

51 - 200

🏗️ Construction

⚡ Energy

SCADA Engineer developing and managing monitoring control systems for renewable energy projects. Engaging with clients and contributing to system designs and improvements.

🕒 June 22

Rune Technologies

11 - 50

📦 Logistics

🏭 Manufacturing

🎖️ Defense

Forward Deployed Engineer at Rune Technologies developing software solutions for military logistics. Collaborating with teams to deliver high-stakes projects and field-test systems.

🕒 June 22

Tern

11 - 50

💼 Consulting

📦 Logistics

📣 Marketing

Implementation Tooling Engineer at Tern, enhancing agency migrations through bulk operations. Focused on backend engineering and data pipeline optimization.