Senior Platform – Site Reliability Engineer (SRE)

Job not on LinkedIn

🔥 14 minutes ago

🇦🇺 Australia – Remote

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

👻 Ghost score 25%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of 3M Consultancy

3M Consultancy

1 - 10 employees

Founded 2018

💼 Consulting

📣 Marketing

🎯 Recruiter

Consulting • Marketing • Recruitment

3M Consultancy is an established Information Technology Staffing Firm dedicated to fulfilling hiring needs for businesses of all sizes. With over 15 years of experience, they pride themselves on sourcing top talent by meticulously understanding client requirements and business culture. Their strategic approach, which combines flexibility and a deep understanding of industry demands, ensures that organizations receive candidates who fit their unique needs and goals.

📋 Description

• Manage AWS production environments across Australia and the United States using Terraform • Improve monitoring, logging, alerting, and observability to identify and resolve issues early • Maintain CI/CD pipelines and improve deployment processes, including automated health checks and rollbacks • Ensure PostgreSQL and Amazon RDS databases remain reliable, secure, and optimized for performance • Manage database backups, restoration procedures, query optimization, and scaling • Develop, document, and regularly test disaster recovery procedures • Monitor infrastructure capacity and prepare systems for increasing customer demand • Collaborate with software development teams to improve application reliability and performance • Manage infrastructure maintenance, server updates, security patches, and vulnerability fixes • Investigate production incidents, identify root causes, and implement long-term solutions • Improve infrastructure automation and operational processes without disrupting ongoing development • Take ownership of a growing cloud-based SaaS environment • Maintain a secure, stable, and scalable AWS environment while supporting product development and business growth

🎯 Requirements

• Minimum 5 years of experience in Platform Engineering, Site Reliability Engineering, DevOps, or Cloud Infrastructure • At least 3 years of hands-on experience managing AWS production environments for SaaS applications • Strong experience with Terraform and Infrastructure as Code (IaC) • Hands-on knowledge of CI/CD pipelines, automated deployments, and rollback procedures • Experience with monitoring, alerting, logging, and observability tools • Strong PostgreSQL and Amazon RDS experience, including performance tuning, backups, restores, and scalability • Experience developing and testing disaster recovery plans, including RTO and RPO • Good understanding of AWS security, networking, access management, and infrastructure management • Experience with capacity planning, performance optimization, and production incident resolution • Ability to independently manage production infrastructure and take full technical ownership • Strong problem-solving and communication skills • Familiarity with SOC 2, ISO 27001, or similar compliance standards • Experience working with managed-service providers or external support teams • Background supporting enterprise SaaS applications, particularly in industrial or supply chain environments • Previous experience as the primary or sole Platform/SRE Engineer • Experience managing cloud infrastructure across multiple regions

🏖️ Benefits

• Remote work arrangement • Full-time employment

Apply Now

Similar Jobs

🕒 September 30

Platform.sh

201 - 500

☁️ SaaS

🛍️ eCommerce

🔌 API

Senior Site Reliability Engineer improving Upsun’s cloud application platform reliability, observability, and automation. Building resilient multi-cloud infrastructure for a global remote software team.

Ansible

AWS

Azure

Cloud

Docker

Google Cloud Platform

Grafana

Kubernetes

Linux

OpenStack

Prometheus

Python

Terraform

Go

🕒 September 29

Climavision

11 - 50

💼 Consulting

📦 Logistics

🤖 Artificial Intelligence

Senior SRE ensuring Kubernetes reliability across Climavision’s weather radar and weather intelligence platform. Automating observability, recovery, high availability, and cost optimization across hybrid infrastructure.

Ansible

Azure

Cloud

Distributed Systems

Grafana

Kubernetes

Node.js

Prometheus

Terraform

🕒 June 30

Red Hat

10,000+ employees

🏢 Enterprise

Customer Site Reliability Engineer managing large-scale systems for cloud services at Red Hat. Focused on enhancing service reliability, customer satisfaction, and technical escalation management.

🗣️🇯🇵 Japanese Required

Ansible

AWS

Azure

Cloud

Distributed Systems

Google Cloud Platform

Kubernetes

Linux

OpenShift

Prometheus

TCP/IP

Terraform

Go

🕒 June 4

Omilia - Conversational Intelligence

201 - 500

💼 Consulting

🛡️ Insurance

✈️ Travel

Senior Site Reliability Engineer maintaining production clusters and developing observability solutions. Collaborate with teams to ensure platform reliability and performance using automation and monitoring tools.

Ansible

AWS

Cloud

Docker

Grafana

Kubernetes

Linux

MySQL

NoSQL

Postgres

Prometheus

Python

RDBMS

Redis

TCP/IP

Terraform

VoIP

Go

🕒 March 28

RevenueCat

51 - 200

💼 Consulting

📣 Marketing

☁️ SaaS

Senior DevOps/DevEx Engineer responsible for building internal development tools at RevenueCat. Collaborating with a global remote team across diverse geographic locations.

AWS

Cloud

Docker

Kubernetes

Python