Site Reliability Engineer

🕒 May 29

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Orion Health

Orion Health

501 - 1000 employees

Founded 1993

đŸ„ Healthcare

đŸ€– Artificial Intelligence

Healthcare ‱ Artificial Intelligence

Orion Health is revolutionising global healthcare so that every individual receives the perfect care for them. It’s leading the world away from traditional broken healthcare systems, towards a future where people are in control of their wellness and able to beat illnesses before they appear. With over 30 years of experience and operations in 13 countries, Orion Health develops technology focused on Population Health Management, Precision Medicine, Health Information Exchange, and more. Their innovations in AI and machine learning aim to improve clinician and patient experiences while delivering better health outcomes.

📋 Description

‱ Design, implement, and maintain reliable, scalable, and secure infrastructure that supports Orion Health's products and services. ‱ Define and monitor Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Service Level Agreements (SLAs) to ensure platform reliability and customer satisfaction. ‱ Build and maintain observability solutions, including monitoring, logging, alerting, and tracing capabilities across cloud environments. ‱ Participate in incident response activities, including troubleshooting, root cause analysis, remediation planning, and post-incident reviews. ‱ Lead initiatives to reduce operational toil through automation, Infrastructure as Code (IaC), and self-service capabilities. ‱ Collaborate closely with software engineering teams to improve application reliability, performance, and operational readiness. ‱ Identify and eliminate reliability bottlenecks through performance tuning, capacity planning, and system optimization. ‱ Support infrastructure and platform upgrades, ensuring minimal disruption and maintaining service availability. ‱ Conduct capacity forecasting and scalability planning to meet future business and customer demands. ‱ Develop operational runbooks, standards, and best practices that improve system resilience and operational efficiency. ‱ Champion reliability engineering principles and foster a culture of continuous improvement across teams. ‱ Contribute to disaster recovery, business continuity, and platform resilience initiatives.

🎯 Requirements

‱ 3+ years of experience in Site Reliability Engineering, Platform Engineering, DevOps, Cloud Operations, or Infrastructure Engineering roles. ‱ Experience supporting and operating production cloud environments. ‱ Strong experience with cloud platforms such as AWS, Azure, or Google Cloud Platform. ‱ Experience implementing Infrastructure as Code (IaC) using tools such as Terraform, Bicep, ARM, or CloudFormation. ‱ Experience with containerisation and orchestration technologies such as Docker and Kubernetes. ‱ Experience building and maintaining monitoring, logging, and observability solutions. ‱ Experience managing production incidents and conducting root cause analysis. ‱ Knowledge of CI/CD pipelines and modern software delivery practices. ‱ Experience with automation and scripting using tools such as PowerShell, Bash, Python, or similar. ‱ Understanding of networking, security, high availability, and disaster recovery principles. ‱ Experience supporting highly available, customer-facing applications and services.

Apply Now

Similar Jobs

🕒 May 28

Enigmatic Smile

11 - 50

đŸ’Œ Consulting

📣 Marketing

☁ SaaS

Senior DevOps Engineer focusing on AWS infrastructure for a security-first fintech scale-up. Collaborating with teams to build robust systems aligned with AWS best practices.

AWS

Cloud

Distributed Systems

Docker

EC2

Linux

Terraform

🕒 May 26

Arbor Education

51 - 200

📚 Education

đŸ€ B2B

Site Reliability Technical Lead driving system design and reliability at Arbor. Mentoring engineers while ensuring scalable, robust, and secure platform performance.

AWS

Cloud

Distributed Systems

Docker

Kubernetes

Microservices

Prometheus

Python

Terraform

Go

🕒 May 22

Flosum

201 - 500

đŸ€ B2B

☁ SaaS

Salesforce DevOps Evangelist promoting Flosum's DevOps and data management platform. Creating engaging content and speaking at major Salesforce events to enhance brand recognition and community involvement.

🕒 May 14

SS&C Technologies

10,000+ employees

đŸ’Œ Consulting

đŸ›Ąïž Insurance

📩 Logistics

Senior Site Reliability Engineer focusing on data platform engineering and internal automation at the leading fintech company SS&C. Design and build reliable data pipelines and automate operational processes.

AWS

Azure

Cloud

ETL

Google Cloud Platform

Python

SQL

Tableau

Terraform

🕒 May 12

Menlo Security Inc.

201 - 500

🔒 Cybersecurity

🏱 Enterprise

Platform Infrastructure Engineer managing cloud-native infrastructure services at Menlo Security. Responsible for designing and maintaining scalable, secure solutions on GCP and AWS using Terraform and Kubernetes.

AWS

Cloud

DNS

Google Cloud Platform

Grafana

Kubernetes

Prometheus

Python

TCP/IP

Terraform