Senior Site Reliability Engineer

Job not on LinkedIn

🔥 6 minutes ago

🇦🇺 Australia – Remote

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

👻 Ghost score 10%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Platform.sh

Platform.sh

201 - 500 employees

Founded 2015

☁️ SaaS

🛍️ eCommerce

🔌 API

SaaS • eCommerce • API

Platform. sh is a collaborative cloud application platform that simplifies full-stack web application development. It enables developers to easily build, deploy, run, scale, and iterate their applications, integrating frontend, backend, APIs, databases, services, and security features without the need to manage infrastructure. The platform emphasizes speed, collaboration, and observability, allowing instant creation of preview environments and efficient development processes.

📋 Description

• Lead the evolution of Upsun’s cloud application platform from traditional cloud operations to a proactive, automation-driven SRE model • Own critical engineering workstreams improving reliability, scalability, and operational efficiency across multi-cloud environments • Partner with engineering, product, and platform teams to embed reliability and performance throughout the software delivery lifecycle • Anticipate architectural bottlenecks and drive infrastructure-as-code practices • Establish observability standards supporting long-term system stability and uptime • Architect monitoring, alerting, and logging with Prometheus, Grafana, and ELK Stack • Establish actionable SLIs/SLOs aligned with core business metrics • Design and implement resilient automated infrastructure and workflows with Terraform and Ansible across AWS, GCP, and Azure • Optimize CI/CD pipeline architectures for fast, secure, zero-downtime releases • Guide high-priority incident triage and lead blameless post-mortems • Implement preventative measures to improve system resiliency • Partner with product and software engineering teams to incorporate SRE practices into product roadmaps • Identify performance bottlenecks and evaluate technologies such as eBPF and container orchestration • Follow a four-week rotation balancing engineering and operations through hands-on troubleshooting and engineering innovation • Participate in on-call one week every 4–5 weeks, from 02:00–10:00 UTC, including a weekend shift

🎯 Requirements

• 5+ years of experience in Site Reliability Engineering, Cloud Operations, or DevOps • Proven experience owning reliability for production platforms at scale • Strong proficiency in Go or Python for custom automation tools, custom controllers, or SRE platform components • Advanced hands-on knowledge of Linux operating system internals, kernel parameters, networking protocols, performance profiling, and system troubleshooting • Deep expertise with AWS, GCP, Azure, or OpenStack • Experience with custom tooling built around cloud SDKs • Experience with declarative infrastructure tools such as Terraform • Proven ability to anticipate operational risks, make architectural trade-offs, and lead technical infrastructure initiatives with minimal guidance • Outstanding cross-functional communication skills and a track record of building alignment and fostering an inclusive engineering culture • Legally authorized to work in Western Australia; visa sponsorship unavailable • Successful background check required • Bonus: experience with custom-built orchestration, edge, storage, and operational tooling • Bonus: experience with Docker and production Kubernetes cluster management or containerized deployment architectures • Bonus: familiarity with PaaS architectures or developer-facing cloud platforms

🏖️ Benefits

• Flexible PTO • Company stock options • Professional development budget • Office equipment budget • Wellness budget • Annual team gatherings • Internet reimbursement • Inclusive parental leave • Remote work travel program • Flexible, open, and inclusive work environment • Accommodations available during the hiring process

Apply Now

Similar Jobs

🔥 8 hours ago

Climavision

11 - 50

💼 Consulting

📦 Logistics

🤖 Artificial Intelligence

Senior SRE ensuring Kubernetes reliability across Climavision’s weather radar and weather intelligence platform. Automating observability, recovery, high availability, and cost optimization across hybrid infrastructure.

Ansible

Azure

Cloud

Distributed Systems

Grafana

Kubernetes

Node.js

Prometheus

Terraform

🕒 June 30

Red Hat

10,000+ employees

🏢 Enterprise

Customer Site Reliability Engineer managing large-scale systems for cloud services at Red Hat. Focused on enhancing service reliability, customer satisfaction, and technical escalation management.

🗣️🇯🇵 Japanese Required

Ansible

AWS

Azure

Cloud

Distributed Systems

Google Cloud Platform

Kubernetes

Linux

OpenShift

Prometheus

TCP/IP

Terraform

Go

🕒 June 4

Omilia - Conversational Intelligence

201 - 500

💼 Consulting

🛡️ Insurance

✈️ Travel

Senior Site Reliability Engineer maintaining production clusters and developing observability solutions. Collaborate with teams to ensure platform reliability and performance using automation and monitoring tools.

Ansible

AWS

Cloud

Docker

Grafana

Kubernetes

Linux

MySQL

NoSQL

Postgres

Prometheus

Python

RDBMS

Redis

TCP/IP

Terraform

VoIP

Go

🕒 March 28

RevenueCat

51 - 200

💼 Consulting

📣 Marketing

☁️ SaaS

Senior DevOps/DevEx Engineer responsible for building internal development tools at RevenueCat. Collaborating with a global remote team across diverse geographic locations.

AWS

Cloud

Docker

Kubernetes

Python