
11 - 50 employees
Founded 2016
π’ Enterprise
π Cybersecurity
π§ Hardware
Enterprise β’ Cybersecurity β’ Hardware
STN Incorporated is an enterprise-grade managed IT and cloud infrastructure provider that delivers secure, audit-ready infrastructure for business-critical systems and demanding AI workloads. STN operates a managed operating model offering private CPU clouds, GPU One AI infrastructure, secure networking and storage, and 24/7 human support with high uptime SLAs. Their services include managed infrastructure and cloud operations, cybersecurity operations and incident response, compliance and risk management (SOC 2 Type II, HIPAA-ready), backup and recovery, and enterprise technology procurement and lifecycle management. STN serves enterprises, high-growth SaaS companies, AI builders and model developers, robotics/physical AI firms, and regulated industries such as healthcare.
π₯ 4 minutes ago
π€ Texas β Remote
β° Full Time
π‘ Mid-level
π Senior
β DevOps & Site Reliability Engineer (SRE)
Improve your chances of getting an interview by checking your resume score before you apply.

11 - 50 employees
Founded 2016
π’ Enterprise
π Cybersecurity
π§ Hardware
Enterprise β’ Cybersecurity β’ Hardware
STN Incorporated is an enterprise-grade managed IT and cloud infrastructure provider that delivers secure, audit-ready infrastructure for business-critical systems and demanding AI workloads. STN operates a managed operating model offering private CPU clouds, GPU One AI infrastructure, secure networking and storage, and 24/7 human support with high uptime SLAs. Their services include managed infrastructure and cloud operations, cybersecurity operations and incident response, compliance and risk management (SOC 2 Type II, HIPAA-ready), backup and recovery, and enterprise technology procurement and lifecycle management. STN serves enterprises, high-growth SaaS companies, AI builders and model developers, robotics/physical AI firms, and regulated industries such as healthcare.
β’ Define and operate Service Level Objectives (SLOs) aligned with customer SLAs β’ Build and maintain the observability stack including metrics, logs, traces, and alerting β’ Lead incident response and chair post-incident reviews β’ Drive automation to reduce toil and improve mean-time-to-recover (MTTR) β’ Author and maintain operational runbooks alongside the NOC β’ Manage on-call rotation, escalation paths, and incident-management tooling β’ Coordinate cross-functionally with NOC, Platform Engineering, and Network Engineering β’ Drive chaos engineering, game days, and reliability testing programs β’ Produce SLA performance reports in coordination with the SLA Manager β’ Mentor junior engineers and contribute to engineering culture
β’ 5+ years in SRE, DevOps, or production engineering roles β’ Strong programming skills in Go, Python, or both β’ Hands-on experience operating Kubernetes-based platforms at scale β’ Deep familiarity with observability tooling (Prometheus, Grafana, Datadog, OpenTelemetry) β’ Strong incident management experience including major-incident command
Apply Nowπ₯ 12 minutes ago
Senior DevOps Engineer at Zafran working on compliance certifications and implementing security controls across infrastructure. Collaborating with teams to enhance security posture and ensure regulatory compliance.
πΊπΈ United States β Remote
β° Full Time
π Senior
β DevOps & Site Reliability Engineer (SRE)
AWS
Kubernetes
Python
Terraform
π₯ 15 minutes ago
Site Reliability Engineer ensuring the stability and resilience of Runpod's distributed platform. Collaborating with engineering teams on reliability frameworks and preventing incidents.
πΊπΈ United States β Remote
π΅ $150k - $200k / year
π° $20M Seed on 2024-06
β° Full Time
π‘ Mid-level
π Senior
β DevOps & Site Reliability Engineer (SRE)
Distributed Systems
Grafana
Linux
Prometheus
π₯ 36 minutes ago
Senior Site Reliability Engineer at Talkiatry, building SRE principles for mental health care. Collaborate with teams to minimize outages and improve reliability for patient services.
πΊπΈ United States β Remote
π΅ $160k - $185k / year
β° Full Time
π Senior
β DevOps & Site Reliability Engineer (SRE)
AWS
Grafana
Prometheus
Python
Terraform
TypeScript
π₯ 41 minutes ago
DevOps Engineer responsible for designing, automating, and maintaining CI/CD pipelines. Focus on cloud infrastructure and improving deployment reliability and security.
πΊπΈ United States β Remote
β° Full Time
π‘ Mid-level
π Senior
β DevOps & Site Reliability Engineer (SRE)
AWS
Cloud
Docker
Google Cloud Platform
Kubernetes
Python
Terraform
π₯ 1 hour ago
Site Reliability Engineer optimizing infrastructure resilience and performance for a leading live streaming platform. Driving enhancement and automation of cloud-based infrastructure with a global network.
πΊπΈ United States β Remote
π΅ $169k - $215k / year
β° Full Time
π Senior
β DevOps & Site Reliability Engineer (SRE)
Ansible
Cloud
Django
Docker
Flask
Java
Kubernetes
Linux
Laravel
Python
Rust
Switching
Terraform
Go