Senior Site Reliability Engineer – Hiring Globally

Job not on LinkedIn

🔥 56 minutes ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Cognativ

Cognativ

11 - 50 employees

Founded 2018

💼 Consulting

🥽 AR/VR

🤖 Artificial Intelligence

Consulting • AR/VR • Artificial Intelligence

Cognativ is an app development agency based in Melbourne, Australia. They combine user experience design, product strategy, and software engineering to deliver mobile and web applications, e-commerce systems, and branded digital experiences. Their specialties include UX and UI design, website and app development (iOS, Android, React Native, React. js, node. js, Python), conversion optimisation, and emerging technologies such as AI/machine learning and AR/VR; they position themselves as quality-focused IT services and consulting partners.

📋 Description

• Own the operational health of computer-vision / AI models and associated services. • Manage reliability and operations, defining SLIs and SLOs for key services. • Ensure observability is trustworthy, improving alert quality and building necessary metrics. • Lead incident response, driving recovery and producing actionable postmortems. • Plan capacity and performance to preempt saturation. • Own business continuity, disaster recovery, and backup strategies. • Engineer software solutions for operational efficiency and reliability. • Govern production change management collaboratively and safely. • Maintain CI/CD pipelines and Terraform configurations for the AWS estate. • Harden security and compliance within the system.

🎯 Requirements

• AWS certification is mandatory. A current AWS Certified DevOps Engineer – Professional or AWS Certified Solutions Architect – Professional is strongly preferred. • 10+ years in Site Reliability Engineering or production operations at scale. We do not expect mastery of every area below on day one. We expect real depth in several and the ability to ramp quickly on the rest. • Demonstrated SLO/error-budget practice. You have defined SLIs and SLOs, run against an error budget, and used it to make real decisions. • Strong production observability skills. Deep with metrics, logs, alarming, and dashboards (CloudWatch, OpenTelemetry, and/or Prometheus/Grafana/Datadog), and able to build the instrumentation when it does not exist. • Proven incident command. You have led incidents, owned an on-call rotation, and written postmortems that changed how a system behaved. • Capacity planning and performance experience across compute, databases, and a messaging or streaming system (Kafka/MSK ideal). • Disaster recovery ownership: backups, replication, failover, and tested RPO/RTO. • Software engineering ability for automation. Comfortable writing Python, Golang, and Bash to build reliability tooling, not just configure off-the-shelf tools. • Expert with Terraform (or equivalent IaC) and strong Linux administration (shell plus Linux GNU utils), comfortable from cloud to bare-metal/edge. • Database operations experience with PostgreSQL (time-series a plus). • A reliability mindset: you instrument before you guess, and you write the runbook. • Nice to have • Operating GPU workloads and serving computer-vision or ML models in production (CUDA, Deep Learning AMIs, inference scaling). • Apache MSK / Kafka and streaming-data operations (Kinesis, Kinesis Video Streams). • AWS IoT Core at scale: device provisioning, certificates, secure tunnelling. • Managing a fleet of edge / on-premise devices (golden images, remote update, systemd). • Operating and modernizing legacy systems (Java 8, Jetty, CentOS). • Chaos engineering / game-day practice, and capacity modeling. • Familiarity with Bazel in a monorepo; Cloudflare, Cognito/Auth0, API Gateway.

Apply Now

Similar Jobs

🔥 1 hour ago

National Trust

10,000+ employees

🤲 Charity

🏨 Hospitality

🛒 Retail

DevSecOps Engineer III at National Digital Trust Company securing digital asset infrastructure. Lead modernization in CI/CD, security controls, and cloud-native systems.

Cloud

ITSM

Kubernetes

SDLC

Terraform

🔥 1 hour ago

PayNearMe

201 - 500

💳 Fintech

☁️ SaaS

🤝 B2B

Site Reliability Engineer at PayNearMe, Inc. responsible for infrastructure management and application reliability. Collaborating with cross-functional teams to enhance system performance and incident response.

🇺🇸 United States – Remote

💵 $180k - $200k / year

🔥 Funding within the last year

💰 $50M Series E - PayNearMe on 2025-09

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

Ansible

AWS

Azure

Chef

Cloud

Docker

EC2

Google Cloud Platform

Grafana

Kubernetes

Prometheus

Puppet

Python

Ruby

Ruby on Rails

Splunk

Terraform

Go

🔥 2 hours ago

Made4net

51 - 200

📦 Logistics

☁️ SaaS

🏢 Enterprise

Cloud Operations Engineer supporting AWS infrastructure for supply chain software solutions. Monitoring systems and ensuring reliability in a global operations team.

Ansible

AWS

Cloud

DNS

EC2

Grafana

ITSM

Linux

Microservices

Oracle

Postgres

Python

Terraform

🔥 2 hours ago

Mercadona

10,000+ employees

🛒 Retail

🍽️ Food & Beverage

DevOps Prime overseeing GCH’s cloud infrastructure while directing vendor resources and ensuring security. Responsible for CI/CD strategies, observability, and incident response across platforms.

AWS

Cloud

Terraform

🔥 2 hours ago

Sycurio

51 - 200

☁️ SaaS

🔐 Security

📋 Compliance

Deployment Engineer for Sycurio solutions deployment and testing. Involves supporting installations and training customer support engineers.

AWS

DNS

EC2

Firewalls

JavaScript

Linux

SOAP

Splunk

TCP/IP

VMware

VoIP