Site Reliability Engineer – Observability

Job not on LinkedIn

🔥 3 hours ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of CluePoints

CluePoints

201 - 500 employees

Founded 2012

🏥 Healthcare

🤖 Artificial Intelligence

☁️ SaaS

💰 Private Equity Round - CluePoints on 2020-08

Healthcare • Artificial Intelligence • SaaS

CluePoints is a clinical-trial focused software company that applies advanced statistics, machine learning, and AI to risk-based quality management (RBQM) and data oversight. Their platform offers products such as RBQM, Site Profile & Oversight Tool (SPOT), Medical & Safety Review (MSR), Intelligent Medical Coding, and Intelligent Query Detection (IQD) to detect anomalies, automate coding and query detection, and help sponsors, CROs, and study teams monitor and de-risk clinical studies across phases. CluePoints positions itself as an AI-powered SaaS provider that turns trial data into actionable insights to improve trial quality and regulatory readiness.

📋 Description

• - Own and improveReal User Monitoring (RUM) for customer-facing applications, including browser performance, client-side errors, user journeys, and frontend service dependencies. • - Partner with frontend, product, and engineering teams to improve visibility into user experience, JavaScript/runtime failures, page performance, and customer-impacting issues. • - Establish and maintain end-to-end observabilityacross frontend, backend, infrastructure, and Kubernetes environments using metrics, logs, traces, dashboards, and alerting. • - Evaluate, implement, and operate managed and self-managed observability solutions, helping guide the evolution of the observability stack.**Support and improve observability tooling such as Sentry, Elastic, Grafana, Prometheus, OpenTelemetry, monitoring agents, and related APM platforms.**Define and maintain SLIs, SLOs, and alerting strategies that improve service reliability, reduce noise, and enable faster detection of production issues.**Lead or support incident detection, alert triage, live production troubleshooting, and service restoration across outage, latency, batch, file transfer, and degradation scenarios, in partnership with Support and Production teams.

🎯 Requirements

• - 5+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or Observability Engineering roles. • - Strong hands-on experience with observability and monitoring platforms, including several of the following:Elastic, Grafana, Prometheus, OpenTelemetry, Sentry, monitoring agents, and managed APM/observability platforms. • - Experience implementing and supporting Real User Monitoring (RUM) and frontend/application observability in production environments. • - Ability to work across frontend, backend, and platform teams to improve telemetry, alerting, and incident diagnosis. • - Experience evaluating or operating managed observability platforms and understanding the trade-offs versus self-managed stacks. • (Nice to have) • - Experience supporting ML, AI, or LLM-backed services in production (RAG, LangSmith, Arize Phoenix, LangChain, LangGraph, Azure OpenAI, OpenAI, or Anthropic APIs).

🏖️ Benefits

• **🇵🇱**** What We Offer – Poland** • - **Comprehensive Health Insurance** (medical, dental, and online consultations, 100% employee coverage) • - **Life Insurance** through UNUM • - **Cafeteria Plan** with flexible monthly credits for wellness, entertainment, and travel • - **MultiSport Card**, co-financed 50/50 • - **Employee Capital Plans (PPK)** with **4% employer contribution** • - A hub-based hybrid model that blends flexibility with purpose — connecting teams through collaboration, learning, and a vibrant social culture.

Apply Now

Similar Jobs

🕒 3 days ago

FinteqHub

51 - 200

💳 Fintech

☁️ SaaS

🤝 B2B

DevSecOps Engineer responsible for application security and CI/CD pipeline hardening at FinteqHub. Collaborate with the engineering and security teams to enhance security processes.

Ansible

AWS

Azure

Cloud

Google Cloud Platform

Kubernetes

Python

Terraform

🕒 3 days ago

LITE TECH, INC.

11 - 50

🏥 Healthcare

🏭 Manufacturing

🔬 Science

Senior Platform Engineer for Lite e-Commerce developing cloud solutions and maintaining observability stack. Collaborating with backend teams and supporting infrastructure for data engineering.

🗣️🇵🇱 Polish Required

Airflow

Cloud

Google Cloud Platform

Grafana

Prometheus

🕒 4 days ago

Arista Networks

1001 - 5000

🏢 Enterprise

📡 Telecommunications

Site Reliability Engineer at Arista Networks focused on operating and enhancing infrastructure for product development teams. Collaborating with an EngProd team to drive scalable and secure systems.

Linux

Python

Shell Scripting

Unix

Go

🕒 5 days ago

Akamai Technologies

5001 - 10000

🔒 Cybersecurity

DevOps Engineer operating systems gathering telemetry from the next-gen cloud computing platform. Collaborating across scrum teams for design, development, and maintenance.

Cloud

Distributed Systems

Docker

Hadoop

Java

Kafka

Kubernetes

Linux

Python

Shell Scripting

Spark

TCP/IP

Terraform

🕒 5 days ago

Valtech

5001 - 10000

💼 Consulting

📣 Marketing

☁️ SaaS

Site Reliability Engineer connecting software development and operations at Valtech. Delivering reliable speed and infrastructure to enhance customer experience.

AWS

Azure

Cloud

Docker

Google Cloud Platform

Java

Kafka

Kubernetes

Microservices

Spring Boot

SpringBoot