Customer Site Reliability Engineer – OpenShift Managed Cloud Services, Spoken Japanese, Kubernetes/AWS/Azure, Linux

🕒 June 30

🇦🇺 Australia – Remote

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

👻 Ghost score 32%

infoinfo

🗣️🇯🇵 Japanese Required

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Red Hat

Red Hat

10,000+ employees

Founded 1993

🏢 Enterprise

💰 Corporate Round on 1999-03

Enterprise • Cloud

Red Hat is a leading provider of enterprise open source software solutions, helping companies worldwide to build and deploy applications across hybrid cloud infrastructures. With a strong focus on developing secure, stable, and innovative technologies, Red Hat offers a broad portfolio including products like Red Hat Enterprise Linux, Red Hat OpenShift, and Red Hat Ansible Automation Platform. These products support IT services on any infrastructure efficiently. Trusted by more than 90% of the U. S. Fortune 500, Red Hat empowers organizations to modernize their IT environments, leveraging open source communities to drive technological advancement.

📋 Description

• Manage large-scale, distributed systems, focusing on minimizing downtime and improving system resilience. • Maintain customer trust and confidence by ensuring stability and functionality of services. • Drive continuous enhancement of processes, tools, and methodologies to support the evolving needs of the service. • Lead the development of code and automation scripts to optimize the scalability, reliability, and performance of services. • Lead and participate in high-priority customer escalations, adopting a customer-first mindset. • Coordinate and execute complex incident response procedures, ensuring timely resolution and thorough postmortems. • Collaborate with cross-functional teams to enhance system robustness. • Demonstrate a proactive mindset to help preempt escalations and ensure reliable operations. • Document resolutions, root causes, and best practices to enrich the knowledge base and promote self-service solutions. • Mentor and coach team members, fostering a culture of continuous learning, knowledge sharing and collaboration. • Participate in on-call rotation and provide leadership during critical incidents. • Collaborate on strategic AI and automation projects designed to increase the efficiency of fleet operations and troubleshooting, ultimately delivering a better product experience for customers.

🎯 Requirements

• Advanced Experience with OpenShift/Kubernetes container platform support or administration. • Proficient with container-based technologies on Linux. • Proficient in managing Linux-based systems in a public cloud such as AWS, Azure, or GCP. • Advanced experience with enterprise systems monitoring; knowledge of Prometheus is preferred. • Advanced with enterprise configuration management such as Ansible, Terraform. • Software engineering experience using object-oriented languages; golang is preferred. • Superior communications skills and experience working directly with and presenting to customers. • Ability to quickly learn new technologies and follow industry trends. • Demonstrated ability to quickly and accurately troubleshoot systems issues. • Solid understanding of standard TCP/IP networking and common protocols. • Fluent in English and any additional language like Japanese, Chinese, Korean, Spanish is an advantage.

🏖️ Benefits

• Flexible working hours • Professional development opportunities

Apply Now

Similar Jobs

🕒 June 4

Omilia - Conversational Intelligence

201 - 500

💼 Consulting

🛡️ Insurance

✈️ Travel

Senior Site Reliability Engineer maintaining production clusters and developing observability solutions. Collaborate with teams to ensure platform reliability and performance using automation and monitoring tools.

Ansible

AWS

Cloud

Docker

Grafana

Kubernetes

Linux

MySQL

NoSQL

Postgres

Prometheus

Python

RDBMS

Redis

TCP/IP

Terraform

VoIP

Go

🕒 April 2

ClickHouse

51 - 200

☁️ SaaS

🏢 Enterprise

🤖 Artificial Intelligence

Database Reliability Engineer driving improvements in performance and reliability for ClickHouse. Collaborating with global teams to optimize operations and enhance service reliability.

AWS

Azure

Cloud

Google Cloud Platform

Python

SQL

🕒 March 28

RevenueCat

51 - 200

💼 Consulting

📣 Marketing

☁️ SaaS

Senior DevOps/DevEx Engineer responsible for building internal development tools at RevenueCat. Collaborating with a global remote team across diverse geographic locations.

AWS

Cloud

Docker

Kubernetes

Python

🕒 March 13

ClickHouse

51 - 200

☁️ SaaS

🏢 Enterprise

🤖 Artificial Intelligence

Senior Site Reliability Engineer at ClickHouse leading reliability initiatives for cloud infrastructure. Collaborating with engineering teams to design and implement scalable, fault-tolerant systems.

Ansible

AWS

Azure

Cloud

Docker

Google Cloud Platform

Kubernetes

Puppet

Python

SQL

Terraform

Go