Site Reliability Engineer – AI Enablement

Job not on LinkedIn

🕒 July 8

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of HealthCatalyst Norway

HealthCatalyst Norway

2 - 10 employees

Founded 2011

🏥 Healthcare

💼 Consulting

🤝 B2B

Healthcare • Consulting • B2B

HealthCatalyst Norway is an initiative (HealthCatalyst AS) that develops Norway as a global innovation and test arena for the health industry, increases health-related exports, and strengthens public–private collaboration. It runs export and accelerator programs that help Norwegian health-tech companies enter the German market by providing advisory services, market guidance, networking, trade-fair visibility (e. g. , DMEA), investor access, and targeted events. The organization collaborates with Innovation Norway, Team Norway, and the Norwegian Directorate of Health to promote e‑health, digital bundling across the Nordics and Germany, and stronger cooperation between healthcare providers and suppliers.

📋 Description

• As a Site Reliability Engineer on the Central AI team, you will help Health Catalyst engineer teams adopt AI responsibly and effectively. • Train and coach engineering teams on how to effectively integrate AI into their development workflows, including the use of AI-assisted coding tools, prompt engineering practices, and agentic development patterns. • Evaluate AI system designs submitted through the Central AI intake process, providing actionable guidance on integration patterns, reliability risks, observability gaps, and alignment with AI governance standards. • Serve as a technical resource for the organization’s AI governance framework — helping teams understand and apply policies around model access, data handling, risk tiers, and responsible AI use in practice. • Partner with engineering teams during the design and implementation phases of AI projects, offering hands-on guidance on LLM integration, RAG pipelines, agentic architectures, and AI service patterns. • Bring an SRE perspective to AI systems — advising teams on observability, SLOs, failure modes, and operational readiness for AI-powered services. • Participate in incident calls as a subject matter expert to provide AI-specific guidance when needed. • Contribute to the development of internal standards, reference architectures, and reusable patterns that make it easier for teams to build AI systems correctly the first time. • Work closely with product managers, data scientists, security, and compliance stakeholders to ensure AI implementations meet organizational, regulatory, and clinical requirements. • Maintain clear documentation of AI architecture patterns, governance guidance, and review decisions to support knowledge sharing and organizational learning. • Stay current with the rapidly evolving AI landscape — LLM capabilities, agentic frameworks, AI safety research, and SRE practices for AI systems — and bring relevant insights back to the team.

🎯 Requirements

• Proven experience solutioning and implementing AI systems in production, including LLM API integration (e.g., Azure AI Foundry, Anthropic Claude) and AI-native application patterns. • Hands-on experience with at least one agentic or RAG framework (e.g., LangChain, LlamaIndex, Semantic Kernel, or similar). • Strong SRE or platform engineering background, with working knowledge of observability, reliability principles, and operational best practices. • Ability to evaluate AI architectures for reliability, security, governance alignment, and operational readiness — and communicate findings clearly to both technical and non-technical audiences. • Experience advising or enabling engineering teams: coaching, conducting reviews, or leading training on AI tooling and best practices. • Familiarity with AI governance concepts, including risk tiering, responsible AI principles, prompt safety, and access control for AI services. • Cloud infrastructure experience with Azure or AWS, including managed AI/ML services. • Familiarity with container-based architectures (Docker, Kubernetes) and CI/CD pipelines. • Strong written and verbal communication skills; able to articulate complex AI concepts to audiences of varying technical background. • Highly collaborative, self-directed, and motivated by helping others succeed with new technology.

🏖️ Benefits

• flexible PTO • professional development stipend • meaningful opportunities for career growth and development

Apply Now

Similar Jobs

🕒 July 8

Cribl

501 - 1000

☁️ SaaS

Senior Site Reliability Engineer unlocking the value of observability data for Cribl. Engaging with teams to improve service delivery and reliability in a remote-first environment.

Ansible

AWS

Azure

Cloud

Grafana

JavaScript

Linux

Node.js

Prometheus

Splunk

Terraform

TypeScript

🕒 July 8

ShorePoint Inc

1 - 10

💼 Consulting

🏥 Healthcare

📦 Logistics

DevSecOps Engineer supporting cloud-based cybersecurity data systems in fast-paced public sector environments. Drive operational excellence through engineering, operating, and monitoring data infrastructure.

Ansible

AWS

Cloud

Cyber Security

Kubernetes

Linux

Python

Terraform

🕒 July 8

Health Catalyst

1001 - 5000

🏥 Healthcare

💼 Consulting

📦 Logistics

Site Reliability Engineer on Central AI team supporting AI systems for healthcare organizations at Health Catalyst. Train teams in AI practices and ensure governance and best practices are followed.

AWS

Azure

Cloud

Docker

Kubernetes

🕒 July 7

Quantiphi

1001 - 5000

💼 Consulting

🏥 Healthcare

📦 Logistics

Sr DevOps Specialist responsible for designing enterprise EKS environments. Work with Fortune 500 clients in a fast-growing AI-focused digital engineering company.

AWS

Kubernetes

Terraform

🕒 July 7

Nametag

11 - 50

🏥 Healthcare

🛡️ Insurance

📦 Logistics

Software Engineer focusing on infrastructure and reliability at Nametag for secure digital identity. Designing scalable systems and tooling to enhance engineering productivity.

🇺🇸 United States – Remote

💵 $120k - $190k / year

💰 Series unknown on 2021-02

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

AWS

Cloud

Microservices

Postgres

Terraform

Go