Site Reliability Engineer – AI Enablement

Vaga não está no LinkedIn

🕒 Julho 8

🇺🇸 Estados Unidos – Remoto (EUA)

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🦅 Patrocina Visto H1B

info

🗣️🇺🇸🇬🇧 Inglês obrigatório

Candidatar-se
Encontrar Vagas Remotas Similares

📊 Verifique sua pontuação de currículo para esta vaga

Melhore suas chances de conseguir uma entrevista verificando sua pontuação de currículo antes de se candidatar.

Logo of Health Catalyst

Health Catalyst

1001 - 5000 funcionários

Fundada em 2008

🏥 Saúde

💼 Consultoria

📦 Logística

Healthcare • Consulting • Logistics

A Health Catalyst é uma fornecedora líder de tecnologia e serviços de dados e análises para organizações de saúde, comprometida em ser o catalisador para uma melhoria massiva, mensurável e informada por dados na área da saúde. A empresa capacita organizações com insights habilitados por IA e soluções de dados abrangentes para impulsionar melhorias escaláveis e mensuráveis nos resultados dos pacientes, na eficiência operacional e no desempenho financeiro. Com foco na gestão da saúde populacional, qualidade clínica e engajamento do paciente, a Health Catalyst visa transformar o setor de saúde por meio de decisões baseadas em dados.

Descrição

• As a Site Reliability Engineer on the Central AI team, you will help Health Catalyst engineer teams adopt AI responsibly and effectively. • Train and coach engineering teams on how to effectively integrate AI into their development workflows, including the use of AI-assisted coding tools, prompt engineering practices, and agentic development patterns. • Evaluate AI system designs submitted through the Central AI intake process, providing actionable guidance on integration patterns, reliability risks, observability gaps, and alignment with AI governance standards. • Serve as a technical resource for the organization’s AI governance framework — helping teams understand and apply policies around model access, data handling, risk tiers, and responsible AI use in practice. • Partner with engineering teams during the design and implementation phases of AI projects, offering hands-on guidance on LLM integration, RAG pipelines, agentic architectures, and AI service patterns. • Bring an SRE perspective to AI systems — advising teams on observability, SLOs, failure modes, and operational readiness for AI-powered services. • Participate in incident calls as a subject matter expert to provide AI-specific guidance when needed. • Contribute to the development of internal standards, reference architectures, and reusable patterns that make it easier for teams to build AI systems correctly the first time. • Work closely with product managers, data scientists, security, and compliance stakeholders to ensure AI implementations meet organizational, regulatory, and clinical requirements. • Maintain clear documentation of AI architecture patterns, governance guidance, and review decisions to support knowledge sharing and organizational learning. • Stay current with the rapidly evolving AI landscape — LLM capabilities, agentic frameworks, AI safety research, and SRE practices for AI systems — and bring relevant insights back to the team.

🎯 Requisitos

• Proven experience solutioning and implementing AI systems in production, including LLM API integration (e.g., Azure AI Foundry, Anthropic Claude) and AI-native application patterns. • Hands-on experience with at least one agentic or RAG framework (e.g., LangChain, LlamaIndex, Semantic Kernel, or similar). • Strong SRE or platform engineering background, with working knowledge of observability, reliability principles, and operational best practices. • Ability to evaluate AI architectures for reliability, security, governance alignment, and operational readiness — and communicate findings clearly to both technical and non-technical audiences. • Experience advising or enabling engineering teams: coaching, conducting reviews, or leading training on AI tooling and best practices. • Familiarity with AI governance concepts, including risk tiering, responsible AI principles, prompt safety, and access control for AI services. • Cloud infrastructure experience with Azure or AWS, including managed AI/ML services. • Familiarity with container-based architectures (Docker, Kubernetes) and CI/CD pipelines. • Strong written and verbal communication skills; able to articulate complex AI concepts to audiences of varying technical background. • Highly collaborative, self-directed, and motivated by helping others succeed with new technology.

🏖️ Benefícios

• Flexible PTO • Professional development stipend • Remote-first work environment

Candidatar-se

Vagas Similares

🕒 Julho 7

Quantiphi

1001 - 5000

💼 Consultoria

🏥 Saúde

📦 Logística

Sr DevOps Specialist responsible for designing enterprise EKS environments. Work with Fortune 500 clients in a fast-growing AI-focused digital engineering company.

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Julho 7

Nametag

11 - 50

🏥 Saúde

🛡️ Seguros

📦 Logística

Software Engineer focusing on infrastructure and reliability at Nametag for secure digital identity. Designing scalable systems and tooling to enhance engineering productivity.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $120.000 - $190.000 / ano

💰 Series unknown em 2021-02

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Julho 7

Flywire

1001 - 5000

🏥 Saúde

✈️ Turismo

📦 Logística

Manager II, Site Reliability Engineering responsible for cloud-based infrastructure performance at Flywire. Collaborating with engineering teams to ensure reliability, automation, and efficiency in production systems.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $160.000 - $200.000 / ano

💰 $60.000.000 Series F em 2021-03

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🦅 Patrocina Visto H1B

info

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Julho 7

Empower

10.000+ funcionários

💸 Finanças

💳 Fintech

👥 B2C

Lead Site Reliability Engineer overseeing reliability initiatives and SRE best practices at Empower. Architecting AWS infrastructure and mentoring SRE teams in a collaborative environment.

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Julho 6

Infarsight

51 - 200

🤖 Inteligência Artificial

✈️ Turismo

📦 Logística

Senior DevOps & Cloud Infrastructure Engineer optimizing AWS environments for automation and product innovation. Leading deployment strategies and resource management across AWS, Vercel, and RackSpace.

🇺🇸 Estados Unidos – Remoto (EUA)

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório