Applied AI Engineer – Systems & Reliability

Job not on LinkedIn

🔥 3 minutes ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of HiPeople

HiPeople

11 - 50 employees

Founded 2020

👥 HR Tech

☁️ SaaS

🤖 Artificial Intelligence

💰 $2.7M Seed Round - HiPeople on 2022-10

HR Tech • SaaS • Artificial Intelligence

HiPeople is an AI-powered hiring platform that helps companies automate and augment recruiting workflows. It offers AI screening to process inbound applications in real time, skills assessments for bias-reduced, skills-based hiring, 24/7 interview capabilities in many languages, and AI-driven, fraud-protected reference checks. The platform integrates deeply with applicant tracking systems and emphasizes privacy and compliance to deploy AI throughout the hiring funnel, reducing recruiter time and accelerating hiring decisions.

📋 Description

• Own evaluation systems and quality standards • Build and maintain evaluation pipelines for core AI workflows across screening, interviews, assessments, and references • Define metrics, benchmarks, and acceptance criteria for AI outputs • Track performance over time (quality trends, drift, regressions) and make results visible across the team • Drive continuous improvement of AI performance • Identify issues across prompts, workflows, and data pipelines using both quantitative analysis and deep dives into real cases • Design and implement improvements across: • prompting strategies • model selection, configuration, and fine-tuning • input data quality and preprocessing • orchestration and workflow design • Push new systems from “working” (80%) to reliable and high-quality (95%+) • Ensure reliability, monitoring, and stability • Build and improve monitoring for AI systems (e.g. dashboards, alerts, tracing) • Detect and prevent failure modes, breakdown risks, and performance degradation • Monitor usage, rate limits, and capacity to ensure stable operation at scale • Drive testing, CI, and safe shipping practices • Integrate AI and prompt testing into CI (e.g. regression tests, golden datasets, staging environments) • Define standards and tooling so product and engineering teams can safely ship without introducing regressions • Act as a quality gate for AI-related changes • Own AI system audits and compliance support • Prepare and support internal and external audits (e.g. SOC 2 and beyond) • Provide evidence, documentation, and artifacts for AI system behavior and controls • Translate audit findings into concrete improvements in systems and processes • Productionize AI workflows (not just prototype them) • Build and productionize AI workflows that meet defined quality and reliability standards • Support product and engineering teams in integrating AI cleanly into product logic and user experience • Ensure new AI capabilities are robust, measurable, and maintainable before release

🎯 Requirements

• 100% alignment with our Ops Principles (if you feel this isn’t you, do not apply) • Excitement for building in Go • Experience working with AI/ML systems, LLMs, or data-intensive applications • High ownership mindset and attention to detail • Strong interest in quality, reliability, and system performance, not just building features • Ability to debug complex systems across prompts, models, and data pipelines • Clear communication and documentation skills • Comfort improving systems and processes, not just using them • Experience with evaluation methods, metrics, or experimentation is a strong plus • Familiarity with monitoring, CI/CD, and production systems is a plus

🏖️ Benefits

• Direct ownership of one of the most critical parts of the company: AI quality and reliability • Work closely with founders on core product and technical decisions • Competitive salary and meaningful stock options • Educational stipend to support ongoing learning and development • The best team to work with (true story!)

Apply Now

Similar Jobs

🔥 6 hours ago

Growe

501 - 1000

🎮 Gaming

🤝 B2B

Building and owning the DevSecOps function in the security team at Growe. Assessing security postures and implementing controls across the software development lifecycle.

🗣️🇺🇦 Ukrainian Required

Ansible

AWS

Azure

Cloud

Google Cloud Platform

Kubernetes

Python

Terraform

🔥 8 hours ago

Ayrin Digital

11 - 50

🤖 Artificial Intelligence

☁️ SaaS

🤝 B2B

DevOps & IT Support Engineer at Ayrin Digital managing multi-cloud environments across AWS, Azure, and GCP. Focus on CI/CD pipelines and IT operations with a strong emphasis on automation.

AWS

Azure

Cloud

DNS

Docker

Firewalls

Google Cloud Platform

Grafana

Jenkins

Kubernetes

Linux

Prometheus

Python

TCP/IP

Terraform

🔥 9 hours ago

Supabase

51 - 200

☁️ SaaS

🔌 API

🤖 Artificial Intelligence

Release Engineer at Supabase, ensuring safe and observable deployments and operational reliability across systems. Engage in incident management, monitoring, and process documentation for improved deployment efficiency.

AWS

Grafana

Kubernetes

Prometheus

Terraform

🕒 6 days ago

Social Discovery Group

1001 - 5000

🌍 Social Impact

📱 Media

DevOps Engineer developing internal services and tools in Go for social discovery products. Collaborating on deployment automation in a remote working environment with a global team.

🗣️🇷🇺 Russian Required

Ansible

Grafana

Kubernetes

Linux

Prometheus

Terraform

Go

🕒 July 11

GitLab

1001 - 5000

💼 Consulting

📣 Marketing

🤖 Artificial Intelligence

Site Reliability Engineer ensuring reliability of GitLab's user-facing services. Supporting operational excellence through engineering principles and automation.

AWS

Cloud

Google Cloud Platform

Kubernetes

Ruby

Terraform

Go