
11 - 50 employees
Founded 2016
đą Enterprise
đ Cybersecurity
đ§ Hardware
Enterprise âą Cybersecurity âą Hardware
STN Incorporated is an enterprise-grade managed IT and cloud infrastructure provider that delivers secure, audit-ready infrastructure for business-critical systems and demanding AI workloads. STN operates a managed operating model offering private CPU clouds, GPU One AI infrastructure, secure networking and storage, and 24/7 human support with high uptime SLAs. Their services include managed infrastructure and cloud operations, cybersecurity operations and incident response, compliance and risk management (SOC 2 Type II, HIPAA-ready), backup and recovery, and enterprise technology procurement and lifecycle management. STN serves enterprises, high-growth SaaS companies, AI builders and model developers, robotics/physical AI firms, and regulated industries such as healthcare.
đ 4 days ago
đ€ Texas â Remote
â° Full Time
đĄ Mid-level
đ Senior
â DevOps & Site Reliability Engineer (SRE)
Improve your chances of getting an interview by checking your resume score before you apply.

11 - 50 employees
Founded 2016
đą Enterprise
đ Cybersecurity
đ§ Hardware
Enterprise âą Cybersecurity âą Hardware
STN Incorporated is an enterprise-grade managed IT and cloud infrastructure provider that delivers secure, audit-ready infrastructure for business-critical systems and demanding AI workloads. STN operates a managed operating model offering private CPU clouds, GPU One AI infrastructure, secure networking and storage, and 24/7 human support with high uptime SLAs. Their services include managed infrastructure and cloud operations, cybersecurity operations and incident response, compliance and risk management (SOC 2 Type II, HIPAA-ready), backup and recovery, and enterprise technology procurement and lifecycle management. STN serves enterprises, high-growth SaaS companies, AI builders and model developers, robotics/physical AI firms, and regulated industries such as healthcare.
âą Define and operate Service Level Objectives (SLOs) aligned with customer SLAs âą Build and maintain the observability stack including metrics, logs, traces, and alerting âą Lead incident response and chair post-incident reviews âą Drive automation to reduce toil and improve mean-time-to-recover (MTTR) âą Author and maintain operational runbooks alongside the NOC âą Manage on-call rotation, escalation paths, and incident-management tooling âą Coordinate cross-functionally with NOC, Platform Engineering, and Network Engineering âą Drive chaos engineering, game days, and reliability testing programs âą Produce SLA performance reports in coordination with the SLA Manager âą Mentor junior engineers and contribute to engineering culture
âą 5+ years in SRE, DevOps, or production engineering roles âą Strong programming skills in Go, Python, or both âą Hands-on experience operating Kubernetes-based platforms at scale âą Deep familiarity with observability tooling (Prometheus, Grafana, Datadog, OpenTelemetry) âą Strong incident management experience including major-incident command
Apply Nowđ 4 days ago
Senior DevOps Engineer at Zafran working on compliance certifications and implementing security controls across infrastructure. Collaborating with teams to enhance security posture and ensure regulatory compliance.
đșđž United States â Remote
â° Full Time
đ Senior
â DevOps & Site Reliability Engineer (SRE)
đ 4 days ago
Site Reliability Engineer ensuring the stability and resilience of Runpod's distributed platform. Collaborating with engineering teams on reliability frameworks and preventing incidents.
đșđž United States â Remote
đ” $150k - $200k / year
đ° $20M Seed on 2024-06
â° Full Time
đĄ Mid-level
đ Senior
â DevOps & Site Reliability Engineer (SRE)
đ 4 days ago
Senior Site Reliability Engineer at Talkiatry, building SRE principles for mental health care. Collaborate with teams to minimize outages and improve reliability for patient services.
đșđž United States â Remote
đ” $160k - $185k / year
â° Full Time
đ Senior
â DevOps & Site Reliability Engineer (SRE)
đ 4 days ago
Site Reliability Engineer optimizing infrastructure resilience and performance for a leading live streaming platform. Driving enhancement and automation of cloud-based infrastructure with a global network.
đșđž United States â Remote
đ” $169k - $215k / year
â° Full Time
đ Senior
â DevOps & Site Reliability Engineer (SRE)
đ 4 days ago
Senior Software Engineer designing and maintaining scalable systems to improve reliability and efficiency at Thumbtack. Collaborating with cross-functional teams to optimize platform services.
đșđž United States â Remote
đ” $179.4k - $232.1k / year
đ° $75M Debt Financing - Thumbtack on 2024-07
â° Full Time
đ Senior
â DevOps & Site Reliability Engineer (SRE)