Site Reliability Engineer

Job not on LinkedIn

🔥 8 minutes ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Storyteller

Storyteller

11 - 50 employees

Founded 2019

☁️ SaaS

📱 Media

🤝 B2B

SaaS • Media • B2B

Storyteller is a developer-focused platform that provides embeddable SDKs, APIs and a CMS to add Stories and TikTok-style vertical video feeds to mobile apps and websites. It offers purpose-built native SDKs for iOS, Android, React Native and Web, plus content management, analytics, native ad support and enterprise features (SSO, SLAs, integrations) so companies can accelerate delivery of immersive short-form content and boost user engagement and revenue.

📋 Description

• Respond to live incidents • Receive automated alerts and technical escalations from Support, then establish customer impact, severity, blast radius and the current system state. • Investigate using logs, metrics, traces, dashboards, deployment history, infrastructure, databases, queues, background jobs, APIs and application code. • Use AI throughout triage and diagnosis while checking its conclusions against real evidence. • Choose and execute a proportionate mitigation, rollback, repair or bounded fix. • Validate that the customer outcome has recovered - not only that an alert has cleared or a dashboard has turned green. • Keep ownership, uncertainty, decisions and next actions visible, and give Support clear technical facts for customer communication. • Join customer conversations occasionally when direct technical involvement is genuinely useful. • Bring in the relevant product team when an incident requires deep product knowledge, a material product decision or a substantial root-cause fix. • Escalate with evidence, customer impact, actions already taken and the specific decision or help required. • Protect developers from routine pages; they should normally be disturbed only for genuine P0/P1 impact or product-specific judgement that cannot safely wait. • Produce a clear incident record and handover, and make sure immediate mitigation, product follow-up and reliability-process follow-up reach the right owners. • Remove, consolidate and tune low-value alerts, and design monitoring around real service and customer outcomes. • Analyse material incidents with AI, validate the conclusions and turn repeated failure patterns into better alerts, runbooks, AI Skills, automation or product improvements. • Improve dashboards, diagnostics, service ownership and escalation information so common incidents are easier to understand and resolve. • Create safe, supervised automation for common operational actions. • Work with product teams to close observability, rollback, runbook and supportability gaps. • Detect and help contain unusual service-cost behaviour, then route wider follow-up to the appropriate cost or product owner. • Make reliability and on-call performance easier for the company to understand and improve over time.

🎯 Requirements

• Previous responsibility for live production systems or an on-call rota is strongly preferred because it is useful evidence that you understand the realities of incident response. • Operational judgement - You can separate customer impact, symptoms and likely causes, make practical decisions under uncertainty and recognise when an intervention is no longer safe or bounded. • Technical comfort and aptitude - You are comfortable exploring unfamiliar systems through code, logs, APIs, data, infrastructure and command-line tools, and can make hands-on changes with a clear validation plan. • AI-native execution - You use AI for substantive technical work - investigation, hypothesis generation, code, automation, incident analysis and workflow improvement - while supervising the agent and challenging its conclusions. • Accuracy and validation discipline - You actively look for false confidence and verify outcomes through appropriate technical and customer signals. • Systems thinking - You look for repeated patterns and improve the triggers, owners, runbooks, automation, metrics and feedback loops around the work. • Clear coordination and communication - You communicate calmly and concisely with Support, developers and non-technical stakeholders, making evidence, impact, uncertainty, ownership and next actions easy to understand. • Curiosity and resilience - You learn unfamiliar products and tools quickly, keep investigating when the first hypothesis fails and change your approach when the evidence demands it. • Nice to have Cloud platforms such as Azure or Cloudflare. • Distributed application and API diagnostics. • Databases, queues and background-processing systems. • Observability, alerting and incident-management platforms. • Infrastructure, deployment and release automation. • Application development and safe production debugging. • AI coding agents and workflow automation.

🏖️ Benefits

• Fully remote working from anywhere in Tunisia!

Apply Now