Site Reliability Engineer

Job not on LinkedIn

🕒 4 days ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of STN Incorporated

STN Incorporated

11 - 50 employees

Founded 2016

🏱 Enterprise

🔒 Cybersecurity

🔧 Hardware

Enterprise ‱ Cybersecurity ‱ Hardware

STN Incorporated is an enterprise-grade managed IT and cloud infrastructure provider that delivers secure, audit-ready infrastructure for business-critical systems and demanding AI workloads. STN operates a managed operating model offering private CPU clouds, GPU One AI infrastructure, secure networking and storage, and 24/7 human support with high uptime SLAs. Their services include managed infrastructure and cloud operations, cybersecurity operations and incident response, compliance and risk management (SOC 2 Type II, HIPAA-ready), backup and recovery, and enterprise technology procurement and lifecycle management. STN serves enterprises, high-growth SaaS companies, AI builders and model developers, robotics/physical AI firms, and regulated industries such as healthcare.

📋 Description

‱ Define and operate Service Level Objectives (SLOs) aligned with customer SLAs ‱ Build and maintain the observability stack including metrics, logs, traces, and alerting ‱ Lead incident response and chair post-incident reviews ‱ Drive automation to reduce toil and improve mean-time-to-recover (MTTR) ‱ Author and maintain operational runbooks alongside the NOC ‱ Manage on-call rotation, escalation paths, and incident-management tooling ‱ Coordinate cross-functionally with NOC, Platform Engineering, and Network Engineering ‱ Drive chaos engineering, game days, and reliability testing programs ‱ Produce SLA performance reports in coordination with the SLA Manager ‱ Mentor junior engineers and contribute to engineering culture

🎯 Requirements

‱ 5+ years in SRE, DevOps, or production engineering roles ‱ Strong programming skills in Go, Python, or both ‱ Hands-on experience operating Kubernetes-based platforms at scale ‱ Deep familiarity with observability tooling (Prometheus, Grafana, Datadog, OpenTelemetry) ‱ Strong incident management experience including major-incident command

Apply Now

Similar Jobs

🕒 4 days ago

Zafran Security

51 - 200

🔐 Security

Senior DevOps Engineer at Zafran working on compliance certifications and implementing security controls across infrastructure. Collaborating with teams to enhance security posture and ensure regulatory compliance.

đŸ‡ș🇾 United States – Remote

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🕒 4 days ago

Runpod

51 - 200

đŸ€– Artificial Intelligence

☁ SaaS

đŸ€ B2B

Site Reliability Engineer ensuring the stability and resilience of Runpod's distributed platform. Collaborating with engineering teams on reliability frameworks and preventing incidents.

đŸ‡ș🇾 United States – Remote

đŸ’” $150k - $200k / year

💰 $20M Seed on 2024-06

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🕒 4 days ago

Talkiatry

501 - 1000

đŸ„ Healthcare

đŸ‘„ B2C

Senior Site Reliability Engineer at Talkiatry, building SRE principles for mental health care. Collaborate with teams to minimize outages and improve reliability for patient services.

đŸ‡ș🇾 United States – Remote

đŸ’” $160k - $185k / year

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🕒 4 days ago

Multi Media, LLC

51 - 200

đŸ’Œ Consulting

📣 Marketing

đŸ“± Media

Site Reliability Engineer optimizing infrastructure resilience and performance for a leading live streaming platform. Driving enhancement and automation of cloud-based infrastructure with a global network.

đŸ‡ș🇾 United States – Remote

đŸ’” $169k - $215k / year

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🕒 4 days ago

Thumbtack

1001 - 5000

đŸȘ Marketplace

☁ SaaS

Senior Software Engineer designing and maintaining scalable systems to improve reliability and efficiency at Thumbtack. Collaborating with cross-functional teams to optimize platform services.

đŸ‡ș🇾 United States – Remote

đŸ’” $179.4k - $232.1k / year

💰 $75M Debt Financing - Thumbtack on 2024-07

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)