
10,000+ employees
Founded 1993
🏢 Enterprise
💰 Corporate Round on 1999-03
Enterprise • Cloud
Red Hat is a leading provider of enterprise open source software solutions, helping companies worldwide to build and deploy applications across hybrid cloud infrastructures. With a strong focus on developing secure, stable, and innovative technologies, Red Hat offers a broad portfolio including products like Red Hat Enterprise Linux, Red Hat OpenShift, and Red Hat Ansible Automation Platform. These products support IT services on any infrastructure efficiently. Trusted by more than 90% of the U. S. Fortune 500, Red Hat empowers organizations to modernize their IT environments, leveraging open source communities to drive technological advancement.
🕒 June 30
🇦🇺 Australia – Remote
⏰ Full Time
🟡 Mid-level
🟠 Senior
⛑ DevOps & Site Reliability Engineer (SRE)
👻 Ghost score 32%
🗣️🇯🇵 Japanese Required
Ansible
AWS
Azure
Cloud
Distributed Systems
Google Cloud Platform
Kubernetes
Linux
OpenShift
Prometheus
TCP/IP
Terraform
Go
Improve your chances of getting an interview by checking your resume score before you apply.

10,000+ employees
Founded 1993
🏢 Enterprise
💰 Corporate Round on 1999-03
Enterprise • Cloud
Red Hat is a leading provider of enterprise open source software solutions, helping companies worldwide to build and deploy applications across hybrid cloud infrastructures. With a strong focus on developing secure, stable, and innovative technologies, Red Hat offers a broad portfolio including products like Red Hat Enterprise Linux, Red Hat OpenShift, and Red Hat Ansible Automation Platform. These products support IT services on any infrastructure efficiently. Trusted by more than 90% of the U. S. Fortune 500, Red Hat empowers organizations to modernize their IT environments, leveraging open source communities to drive technological advancement.
• Manage large-scale, distributed systems, focusing on minimizing downtime and improving system resilience. • Maintain customer trust and confidence by ensuring stability and functionality of services. • Drive continuous enhancement of processes, tools, and methodologies to support the evolving needs of the service. • Lead the development of code and automation scripts to optimize the scalability, reliability, and performance of services. • Lead and participate in high-priority customer escalations, adopting a customer-first mindset. • Coordinate and execute complex incident response procedures, ensuring timely resolution and thorough postmortems. • Collaborate with cross-functional teams to enhance system robustness. • Demonstrate a proactive mindset to help preempt escalations and ensure reliable operations. • Document resolutions, root causes, and best practices to enrich the knowledge base and promote self-service solutions. • Mentor and coach team members, fostering a culture of continuous learning, knowledge sharing and collaboration. • Participate in on-call rotation and provide leadership during critical incidents. • Collaborate on strategic AI and automation projects designed to increase the efficiency of fleet operations and troubleshooting, ultimately delivering a better product experience for customers.
• Advanced Experience with OpenShift/Kubernetes container platform support or administration. • Proficient with container-based technologies on Linux. • Proficient in managing Linux-based systems in a public cloud such as AWS, Azure, or GCP. • Advanced experience with enterprise systems monitoring; knowledge of Prometheus is preferred. • Advanced with enterprise configuration management such as Ansible, Terraform. • Software engineering experience using object-oriented languages; golang is preferred. • Superior communications skills and experience working directly with and presenting to customers. • Ability to quickly learn new technologies and follow industry trends. • Demonstrated ability to quickly and accurately troubleshoot systems issues. • Solid understanding of standard TCP/IP networking and common protocols. • Fluent in English and any additional language like Japanese, Chinese, Korean, Spanish is an advantage.
• Flexible working hours • Professional development opportunities
Apply Now🕒 June 4
Senior Site Reliability Engineer maintaining production clusters and developing observability solutions. Collaborate with teams to ensure platform reliability and performance using automation and monitoring tools.
Ansible
AWS
Cloud
Docker
Grafana
Kubernetes
Linux
MySQL
NoSQL
Postgres
Prometheus
Python
RDBMS
Redis
TCP/IP
Terraform
VoIP
Go
🕒 April 2
Database Reliability Engineer driving improvements in performance and reliability for ClickHouse. Collaborating with global teams to optimize operations and enhance service reliability.
AWS
Azure
Cloud
Google Cloud Platform
Python
SQL
🕒 March 28
Senior DevOps/DevEx Engineer responsible for building internal development tools at RevenueCat. Collaborating with a global remote team across diverse geographic locations.
🇦🇺 Australia – Remote
💵 $227k / year
💰 $40M Series B on 2021-05
⏰ Full Time
🟠 Senior
⛑ DevOps & Site Reliability Engineer (SRE)
AWS
Cloud
Docker
Kubernetes
Python
🕒 March 13
Senior Site Reliability Engineer at ClickHouse leading reliability initiatives for cloud infrastructure. Collaborating with engineering teams to design and implement scalable, fault-tolerant systems.
Ansible
AWS
Azure
Cloud
Docker
Google Cloud Platform
Kubernetes
Puppet
Python
SQL
Terraform
Go