
201 - 500 employees
Founded 2015
☁️ SaaS
🛍️ eCommerce
🔌 API
SaaS • eCommerce • API
Platform. sh is a collaborative cloud application platform that simplifies full-stack web application development. It enables developers to easily build, deploy, run, scale, and iterate their applications, integrating frontend, backend, APIs, databases, services, and security features without the need to manage infrastructure. The platform emphasizes speed, collaboration, and observability, allowing instant creation of preview environments and efficient development processes.
🔥 6 minutes ago
Ansible
AWS
Azure
Cloud
Docker
Google Cloud Platform
Grafana
Kubernetes
Linux
OpenStack
Prometheus
Python
Terraform
Go
Improve your chances of getting an interview by checking your resume score before you apply.

201 - 500 employees
Founded 2015
☁️ SaaS
🛍️ eCommerce
🔌 API
SaaS • eCommerce • API
Platform. sh is a collaborative cloud application platform that simplifies full-stack web application development. It enables developers to easily build, deploy, run, scale, and iterate their applications, integrating frontend, backend, APIs, databases, services, and security features without the need to manage infrastructure. The platform emphasizes speed, collaboration, and observability, allowing instant creation of preview environments and efficient development processes.
• Lead the evolution of Upsun’s cloud application platform from traditional cloud operations to a proactive, automation-driven SRE model • Own critical engineering workstreams improving reliability, scalability, and operational efficiency across multi-cloud environments • Partner with engineering, product, and platform teams to embed reliability and performance throughout the software delivery lifecycle • Anticipate architectural bottlenecks and drive infrastructure-as-code practices • Establish observability standards supporting long-term system stability and uptime • Architect monitoring, alerting, and logging with Prometheus, Grafana, and ELK Stack • Establish actionable SLIs/SLOs aligned with core business metrics • Design and implement resilient automated infrastructure and workflows with Terraform and Ansible across AWS, GCP, and Azure • Optimize CI/CD pipeline architectures for fast, secure, zero-downtime releases • Guide high-priority incident triage and lead blameless post-mortems • Implement preventative measures to improve system resiliency • Partner with product and software engineering teams to incorporate SRE practices into product roadmaps • Identify performance bottlenecks and evaluate technologies such as eBPF and container orchestration • Follow a four-week rotation balancing engineering and operations through hands-on troubleshooting and engineering innovation • Participate in on-call one week every 4–5 weeks, from 02:00–10:00 UTC, including a weekend shift
• 5+ years of experience in Site Reliability Engineering, Cloud Operations, or DevOps • Proven experience owning reliability for production platforms at scale • Strong proficiency in Go or Python for custom automation tools, custom controllers, or SRE platform components • Advanced hands-on knowledge of Linux operating system internals, kernel parameters, networking protocols, performance profiling, and system troubleshooting • Deep expertise with AWS, GCP, Azure, or OpenStack • Experience with custom tooling built around cloud SDKs • Experience with declarative infrastructure tools such as Terraform • Proven ability to anticipate operational risks, make architectural trade-offs, and lead technical infrastructure initiatives with minimal guidance • Outstanding cross-functional communication skills and a track record of building alignment and fostering an inclusive engineering culture • Legally authorized to work in Western Australia; visa sponsorship unavailable • Successful background check required • Bonus: experience with custom-built orchestration, edge, storage, and operational tooling • Bonus: experience with Docker and production Kubernetes cluster management or containerized deployment architectures • Bonus: familiarity with PaaS architectures or developer-facing cloud platforms
• Flexible PTO • Company stock options • Professional development budget • Office equipment budget • Wellness budget • Annual team gatherings • Internet reimbursement • Inclusive parental leave • Remote work travel program • Flexible, open, and inclusive work environment • Accommodations available during the hiring process
Apply Now🔥 8 hours ago
Senior SRE ensuring Kubernetes reliability across Climavision’s weather radar and weather intelligence platform. Automating observability, recovery, high availability, and cost optimization across hybrid infrastructure.
🇦🇺 Australia – Remote
💵 $130k - $170k / year
💰 $100M Series A on 2021-06
⏰ Full Time
🟠 Senior
⛑ DevOps & Site Reliability Engineer (SRE)
Ansible
Azure
Cloud
Distributed Systems
Grafana
Kubernetes
Node.js
Prometheus
Terraform
🕒 June 30
Customer Site Reliability Engineer managing large-scale systems for cloud services at Red Hat. Focused on enhancing service reliability, customer satisfaction, and technical escalation management.
🇦🇺 Australia – Remote
💰 Corporate Round on 1999-03
⏰ Full Time
🟡 Mid-level
🟠 Senior
⛑ DevOps & Site Reliability Engineer (SRE)
🗣️🇯🇵 Japanese Required
Ansible
AWS
Azure
Cloud
Distributed Systems
Google Cloud Platform
Kubernetes
Linux
OpenShift
Prometheus
TCP/IP
Terraform
Go
🕒 June 4
Senior Site Reliability Engineer maintaining production clusters and developing observability solutions. Collaborate with teams to ensure platform reliability and performance using automation and monitoring tools.
Ansible
AWS
Cloud
Docker
Grafana
Kubernetes
Linux
MySQL
NoSQL
Postgres
Prometheus
Python
RDBMS
Redis
TCP/IP
Terraform
VoIP
Go
🕒 March 28
Senior DevOps/DevEx Engineer responsible for building internal development tools at RevenueCat. Collaborating with a global remote team across diverse geographic locations.
🇦🇺 Australia – Remote
💵 $227k / year
💰 $40M Series B on 2021-05
⏰ Full Time
🟠 Senior
⛑ DevOps & Site Reliability Engineer (SRE)
AWS
Cloud
Docker
Kubernetes
Python