Senior Software Engineer – SRE, Observability Tooling

🕒 July 15

🇪🇪 Estonia – Remote

⏰ Full Time

🟠 Senior

🧑‍💻 Full-stack Engineer

👻 Ghost score 11%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Glia

Glia

201 - 500 employees

Founded 2012

💼 Consulting

📦 Logistics

🤖 Artificial Intelligence

💰 $45M Series D - Glia on 2022-03

Consulting • Logistics • Artificial Intelligence

Glia is an AI-powered customer service platform and ChannelLess® contact center solution that helps banks, credit unions, and financial institutions automate and personalize voice and digital interactions. The company provides virtual assistants, voice AI, agent co-pilots, manager AI analysts, predictive routing, cognitive quality management, and analytics to boost self-service, agent productivity, compliance, and growth metrics (loans, deposits, cost savings). Glia emphasizes security, integrations, and industry-specific features for banking and credit unions, delivering a SaaS platform that replaces legacy CCaaS systems.

📋 Description

• Focus on building SRE and observability tooling — the platform, automation, and standards other teams use to keep their services healthy. • Developing standards, infrastructure and automation for dashboards, alerts, and monitors as code. • Partnering with development teams to establish production readiness and operational readiness. • Building the tooling and templates teams use to define, measure, and report on Service Level Objectives (SLOs) and Service Level Indicators (SLIs) for their services. • Developing tooling to automate observability and operational workflows, eliminating manual toil for engineering teams. • Building and improving the incident response tooling and workflows that help teams resolve outages faster and learn from them.

🎯 Requirements

• Expert-level proficiency with AWS and Kubernetes (EKS), particularly in areas of observability, networking, and auto-scaling. • Experience with modern observability platforms (e.g., DataDog, Prometheus) and a deep understanding of metrics, logging, and tracing. • Deep, practical understanding of Site Reliability Engineering (SRE) principles (SLOs, error budgets, toil reduction). • Demonstrable experience analyzing and troubleshooting large-scale distributed systems. • Strong software development skills in a language like Python or Go, used to build operational tools, services, or automation. • Expertise in designing and operating robust CI/CD pipelines for a microservices architecture (e.g., using ArgoCD, Github Actions, Helm). • A systematic, data-driven approach to problem-solving and root cause analysis. • Proficiency in using AI tools thoughtfully, maintaining ownership of the final output while recognizing the tools' limitations.

🏖️ Benefits

• Flexible remote collaboration • Optional offices in Tallinn and Tartu • In-person innovation and connection sessions twice a year

Apply Now

Similar Jobs

🕒 June 26

Veriff

501 - 1000

🔐 Security

📋 Compliance

💸 Finance

Distributed Systems

Python

🕒 April 1

Modash

11 - 50

👥 B2C

📣 Marketing

☁️ SaaS

Senior Product Engineer at Modash working on a creator discovery and management platform for Instagram, TikTok, and YouTube. Join a growing team to build meaningful software solutions.

Cloud