Site Reliability Engineer

Job not on LinkedIn

🕒 July 28

🤠 Texas – Remote

infoinfo

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

👻 Ghost score 18%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of STN Incorporated

STN Incorporated

11 - 50 employees

Founded 2016

🏢 Enterprise

🔒 Cybersecurity

🔧 Hardware

Enterprise • Cybersecurity • Hardware

STN Incorporated is an enterprise-grade managed IT and cloud infrastructure provider that delivers secure, audit-ready infrastructure for business-critical systems and demanding AI workloads. STN operates a managed operating model offering private CPU clouds, GPU One AI infrastructure, secure networking and storage, and 24/7 human support with high uptime SLAs. Their services include managed infrastructure and cloud operations, cybersecurity operations and incident response, compliance and risk management (SOC 2 Type II, HIPAA-ready), backup and recovery, and enterprise technology procurement and lifecycle management. STN serves enterprises, high-growth SaaS companies, AI builders and model developers, robotics/physical AI firms, and regulated industries such as healthcare.

📋 Description

• Define and operate Service Level Objectives (SLOs) aligned with customer SLAs • Build and maintain the observability stack including metrics, logs, traces, and alerting • Lead incident response and chair post-incident reviews • Drive automation to reduce toil and improve mean-time-to-recover (MTTR) • Author and maintain operational runbooks alongside the NOC • Manage on-call rotation, escalation paths, and incident-management tooling • Coordinate cross-functionally with NOC, Platform Engineering, and Network Engineering • Drive chaos engineering, game days, and reliability testing programs • Produce SLA performance reports in coordination with the SLA Manager • Mentor junior engineers and contribute to engineering culture

🎯 Requirements

• 5+ years in SRE, DevOps, or production engineering roles • Strong programming skills in Go, Python, or both • Hands-on experience operating Kubernetes-based platforms at scale • Deep familiarity with observability tooling (Prometheus, Grafana, Datadog, OpenTelemetry) • Strong incident management experience including major-incident command

Apply Now

Similar Jobs

🕒 July 28

Runpod

51 - 200

🤖 Artificial Intelligence

☁️ SaaS

🤝 B2B

Site Reliability Engineer ensuring the stability and resilience of Runpod's distributed platform. Collaborating with engineering teams on reliability frameworks and preventing incidents.

🕒 July 28

Multi Media, LLC

51 - 200

💼 Consulting

📣 Marketing

📱 Media

Site Reliability Engineer optimizing infrastructure resilience and performance for a leading live streaming platform. Driving enhancement and automation of cloud-based infrastructure with a global network.

🕒 July 28

Thumbtack

1001 - 5000

🏪 Marketplace

☁️ SaaS

Senior Software Engineer designing and maintaining scalable systems to improve reliability and efficiency at Thumbtack. Collaborating with cross-functional teams to optimize platform services.

🇺🇸 United States – Remote

💵 $179.4k - $232.1k / year

💰 $75M Debt Financing - Thumbtack on 2024-07

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🕒 July 28

ICF

5001 - 10000

💼 Consulting

🏛️ Government

🏥 Healthcare

DevOps Engineer building healthcare reporting services for ICF. Implementing cloud-based solutions using AWS and fostering collaboration on CI/CD pipeline improvements.

🕒 July 28

ICF

5001 - 10000

💼 Consulting

🏛️ Government

🏥 Healthcare

Senior DevOps Engineer delivering best in class healthcare reporting services for ICF. Working collaboratively to implement cloud solutions and establish CI/CD pipelines using AWS.