Senior Site Reliability Engineer, AUS

Job not on LinkedIn

🔥 0 minutes ago

🇦🇺 Australia – Remote

💵 $130k - $170k / year

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

👻 Ghost score 0%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Climavision

Climavision

11 - 50 employees

💼 Consulting

📦 Logistics

🤖 Artificial Intelligence

💰 $100M Series A on 2021-06

Consulting • Logistics • Artificial Intelligence

Climavision is a pioneering weather intelligence company that revolutionizes weather forecasting through its advanced suite of AI-powered products and services. With a focus on precision and customization, Climavision offers hyper-accurate weather insights for various industries, including agriculture, energy, and transportation. Their innovative Horizon AI models integrate diverse data sources, enabling businesses and government agencies to make informed decisions by anticipating weather-related risks and optimizing operations.

📋 Description

• Own production reliability for customer-facing platform and data services across Azure, colocation, and edge Kubernetes environments. • Support the whole company through a shared SRE function across radar network and weather intelligence operations. • Define and improve SLIs, SLOs, alerting standards, and operational metrics. • Build shared observability dashboards and alerting for the full fleet. • Design and build automated recovery and self-healing for production systems. • Optimize cluster resources and costs, including right-sizing workloads and nodes and migrating workloads off Azure. • Coordinate production incident response, including troubleshooting, mitigation, communication, and postmortem analysis. • Diagnose and resolve complex issues across application services, Kubernetes infrastructure, storage, and distributed systems. • Drive multi-replica and multi-cluster high availability, including workload placement, scheduling, failover, traffic routing, and data replication. • Operate and improve self-managed Kubernetes across cloud-hosted, colocation, and edge clusters. • Perform Kubernetes upgrades, patching, cluster health management, node management, and production change management. • Improve observability, autoscaling, ingress, distributed storage, resiliency, and operational maturity. • Design and validate Kubernetes workloads for resiliency, scalability, and operational efficiency. • Partner with software engineering teams on production readiness, deployment safety, resiliency, and operational visibility. • Maintain deployment pipelines, Helm charts, Kubernetes manifests, and infrastructure automation. • Support metrics, logging, distributed tracing, dashboarding, and alerting platforms. • Conduct performance engineering and capacity planning for peak weather-event demand. • Facilitate blameless postmortems and complete operational follow-up actions. • Improve disaster recovery, failover, and business continuity across cloud, colocation, and edge environments. • Drive automation, toil reduction, game days, production readiness reviews, and reliability best practices. • Serve as a senior technical resource and mentor. • Participate in separate rotating weekday and weekend on-call schedules, approximately every five weeks each.

🎯 Requirements

• A bachelor's degree in computer science, software engineering, or a related field; equivalent professional experience considered. • Minimum of 7 years of experience in Site Reliability Engineering, DevOps, Production Engineering, Platform Engineering, or a related infrastructure-focused role, with at least 4 years in a role formally titled Site Reliability Engineer or carrying explicit SLO / error-budget accountability. • Deep, hands-on experience operating native Kubernetes; managed distributions such as AKS and EKS are acceptable. • Experience optimizing Kubernetes clusters, including right-sizing workloads and node pools, resource management, and reducing infrastructure cost. • Experience building dashboards, metrics pipelines, and alerting. • Experience designing and operating workloads for safe horizontal scaling across multiple replicas, including idempotency, concurrency, and state handling. • Experience designing or operating multi-cluster high-availability architectures, including failover behavior, traffic routing, and cross-cluster service deployment. • Experience supporting customer-facing production systems with uptime, reliability, and incident-response responsibilities. • Experience diagnosing and resolving production incidents across application, platform, and Kubernetes infrastructure layers. • Experience operating Kubernetes outside strictly managed cloud environments, including bare-metal, colocation, edge, or hybrid infrastructure. • Experience with Kubernetes tooling and ecosystem technologies such as Rancher, Helm, autoscaling frameworks, observability stacks, or distributed storage systems. • Strong understanding of infrastructure automation and Infrastructure as Code using tools such as Terraform and Ansible. • Experience supporting CI/CD and production deployment pipelines; GitHub Actions is used at Climavision. • Experience with monitoring, logging, and observability platforms such as DataDog, Prometheus, Grafana, Loki, OpenTelemetry, or comparable technologies. • Experience operating distributed systems and microservice-based architectures in production. • Working knowledge of Microsoft Azure infrastructure. • Strong troubleshooting skills across infrastructure, application, and platform layers. • Experience participating in a structured production on-call rotation supporting business-critical systems. • Working familiarity with Jira, Confluence, and Microsoft Entra. • Strong written and verbal communication skills, including incident documentation and postmortem authoring. • Experience working in start-up, scale-up, or other fast-moving engineering environments. • Any offer of employment is contingent on completion of a background check to company standard.

🏖️ Benefits

• Benefits of a dynamic and growing organization • A challenging, hands-on role that will have real impact on the business • Competitive compensation • Comprehensive benefits package • 401(k) Savings Plan • Medical/Dental/Vision Benefits • Health Savings Account (HSA) and Flexible Spending Account (FSA) • Unlimited Paid Time-off • 11 Paid Holidays • Paid Parental Leave • Company Paid Short-term Disability (STD) • Company Paid Long-term Disability (LTD) • Company Paid Life Insurance • Rotating weekday and weekend on-call schedule with separate rotations

Apply Now

Similar Jobs

🕒 June 30

Red Hat

10,000+ employees

🏢 Enterprise

Customer Site Reliability Engineer managing large-scale systems for cloud services at Red Hat. Focused on enhancing service reliability, customer satisfaction, and technical escalation management.

🗣️🇯🇵 Japanese Required

Ansible

AWS

Azure

Cloud

Distributed Systems

Google Cloud Platform

Kubernetes

Linux

OpenShift

Prometheus

TCP/IP

Terraform

Go

🕒 June 4

Omilia - Conversational Intelligence

201 - 500

💼 Consulting

🛡️ Insurance

✈️ Travel

Senior Site Reliability Engineer maintaining production clusters and developing observability solutions. Collaborate with teams to ensure platform reliability and performance using automation and monitoring tools.

Ansible

AWS

Cloud

Docker

Grafana

Kubernetes

Linux

MySQL

NoSQL

Postgres

Prometheus

Python

RDBMS

Redis

TCP/IP

Terraform

VoIP

Go

🕒 March 28

RevenueCat

51 - 200

💼 Consulting

📣 Marketing

☁️ SaaS

Senior DevOps/DevEx Engineer responsible for building internal development tools at RevenueCat. Collaborating with a global remote team across diverse geographic locations.

AWS

Cloud

Docker

Kubernetes

Python