Site Reliability Engineer

Job not on LinkedIn

πŸ”₯ 4 minutes ago

Apply Now
Find Similar Remote Jobs

πŸ“Š Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of STN Incorporated

STN Incorporated

11 - 50 employees

Founded 2016

🏒 Enterprise

πŸ”’ Cybersecurity

πŸ”§ Hardware

Enterprise β€’ Cybersecurity β€’ Hardware

STN Incorporated is an enterprise-grade managed IT and cloud infrastructure provider that delivers secure, audit-ready infrastructure for business-critical systems and demanding AI workloads. STN operates a managed operating model offering private CPU clouds, GPU One AI infrastructure, secure networking and storage, and 24/7 human support with high uptime SLAs. Their services include managed infrastructure and cloud operations, cybersecurity operations and incident response, compliance and risk management (SOC 2 Type II, HIPAA-ready), backup and recovery, and enterprise technology procurement and lifecycle management. STN serves enterprises, high-growth SaaS companies, AI builders and model developers, robotics/physical AI firms, and regulated industries such as healthcare.

πŸ“‹ Description

β€’ Define and operate Service Level Objectives (SLOs) aligned with customer SLAs β€’ Build and maintain the observability stack including metrics, logs, traces, and alerting β€’ Lead incident response and chair post-incident reviews β€’ Drive automation to reduce toil and improve mean-time-to-recover (MTTR) β€’ Author and maintain operational runbooks alongside the NOC β€’ Manage on-call rotation, escalation paths, and incident-management tooling β€’ Coordinate cross-functionally with NOC, Platform Engineering, and Network Engineering β€’ Drive chaos engineering, game days, and reliability testing programs β€’ Produce SLA performance reports in coordination with the SLA Manager β€’ Mentor junior engineers and contribute to engineering culture

🎯 Requirements

β€’ 5+ years in SRE, DevOps, or production engineering roles β€’ Strong programming skills in Go, Python, or both β€’ Hands-on experience operating Kubernetes-based platforms at scale β€’ Deep familiarity with observability tooling (Prometheus, Grafana, Datadog, OpenTelemetry) β€’ Strong incident management experience including major-incident command

Apply Now

Similar Jobs

πŸ”₯ 12 minutes ago

Zafran Security

51 - 200

πŸ” Security

Senior DevOps Engineer at Zafran working on compliance certifications and implementing security controls across infrastructure. Collaborating with teams to enhance security posture and ensure regulatory compliance.

AWS

Kubernetes

Python

Terraform

πŸ”₯ 15 minutes ago

Runpod

51 - 200

πŸ€– Artificial Intelligence

☁️ SaaS

🀝 B2B

Site Reliability Engineer ensuring the stability and resilience of Runpod's distributed platform. Collaborating with engineering teams on reliability frameworks and preventing incidents.

Distributed Systems

Grafana

Linux

Prometheus

πŸ”₯ 36 minutes ago

Talkiatry

501 - 1000

πŸ₯ Healthcare

πŸ‘₯ B2C

Senior Site Reliability Engineer at Talkiatry, building SRE principles for mental health care. Collaborate with teams to minimize outages and improve reliability for patient services.

AWS

Grafana

Prometheus

Python

Terraform

TypeScript

πŸ”₯ 41 minutes ago

Careerswift

2 - 10

πŸ‘₯ HR Tech

🎯 Recruiter

☁️ SaaS

DevOps Engineer responsible for designing, automating, and maintaining CI/CD pipelines. Focus on cloud infrastructure and improving deployment reliability and security.

AWS

Cloud

Docker

Google Cloud Platform

Kubernetes

Python

Terraform

πŸ”₯ 1 hour ago

Multi Media, LLC

51 - 200

πŸ’Ό Consulting

πŸ“£ Marketing

πŸ“± Media

Site Reliability Engineer optimizing infrastructure resilience and performance for a leading live streaming platform. Driving enhancement and automation of cloud-based infrastructure with a global network.

Ansible

Cloud

Django

Docker

Flask

Java

Kubernetes

Linux

Laravel

Python

Rust

Switching

Terraform

Go