Site Reliability Engineer

Job not on LinkedIn

🔥 0 minutes ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Climb Channel Solutions NA

Climb Channel Solutions NA

51 - 200 employees

Founded 1982

💼 Consulting

📦 Logistics

📣 Marketing

Consulting • Logistics • Marketing

Climb Channel Solutions NA is an IT distribution company focused on providing leading and innovative technology solutions. They support technology resellers by offering expertise in areas such as virtualization, cloud, data management, and cybersecurity, thereby enhancing their partners' success. Climb is dedicated to transforming IT distribution with exceptional service and an extensive vendor marketplace to facilitate business growth for its partners across various sectors, including public and private markets.

📋 Description

• Own the availability and performance of production SaaS applications running on Azure across multiple geographic regions • Lead troubleshooting and resolution of cloud infrastructure and application issues, including AKS failures, deployment rollbacks, ingress and networking issues, and autoscaling problems • Participate in an on-call rotation, including weekends, and drive incident response from detection through resolution • Improve disaster recovery, failover, and incident management processes across multi-region deployments • Build and maintain automation scripts and monitoring tools to reduce manual toil • Author post-incident reviews and root-cause analyses, and drive preventive actions to closure • Partner with senior engineers and cross-functional teams on reliability, observability, and performance best practices • Contribute to continuous improvement across infrastructure, tooling, and processes • Communicate with customer-facing stakeholders during incidents through status updates and written incident summaries

🎯 Requirements

• 5+ years of relevant experience in Site Reliability Engineering, DevOps, or Cloud Administration, with ownership of production systems • Hands-on experience administering Azure environments, including AKS/Kubernetes, core Azure services, cloud networking, and cloud security fundamentals • Solid understanding of monitoring, logging, and alerting practices, including Datadog, Azure Monitor, or ELK stack • Hands-on troubleshooting with log analysis and stack traces using Datadog APM • Familiarity with firewalls, load balancers, VPNs, DNS, and routing • Experience with automation and scripting using PowerShell, Python, or similar • Understanding of cloud backup, redundancy, and disaster recovery strategies, including geo-redundant and multi-region deployments • Ownership across the full incident lifecycle, from detection through post-mortem • Clear and professional written communication for incident updates and summaries • Experience with AWS Cloud Platform preferred • Experience with CI/CD tools such as Azure DevOps preferred • Experience with infrastructure-as-code tools such as Terraform or ARM templates preferred • Prior experience operating SaaS products with regional tenant architectures preferred • Successful candidates must complete a comprehensive criminal background check, education verification, and employment verification

🏖️ Benefits

• Competitive salaries • Meaningful bonus program • Healthcare insurance • Pension/retirement matching • Comprehensive life insurance • Employee assistance program • Time off plans • Paid company holidays • Career progression • Culture of innovation • Global, collaborative work environment

Apply Now

Similar Jobs

🔥 2 hours ago

Salve.Inno

11 - 50

💼 Consulting

📣 Marketing

📦 Logistics

Senior Site Reliability Engineer operating Kubernetes, AWS, and observability platforms for mission-critical cloud services. Automating operations, strengthening incident response, and improving reliability for customers worldwide.

Ansible

AWS

DNS

Grafana

Kubernetes

Linux

MySQL

NoSQL

Postgres

Prometheus

Python

Redis

TCP/IP

Terraform

VoIP

Go

🕒 Yesterday

DysrupIT

51 - 200

🏢 Enterprise

☁️ SaaS

🔒 Cybersecurity

DevOps Engineer building and supporting scalable AWS and Nutanix infrastructure for DysrupIT’s technology consulting clients. Automating deployments, monitoring platforms, and resolving incidents across enterprise environments.

Ansible

AWS

Cloud

DNS

Docker

Kubernetes

Linux

Microservices

MongoDB

MySQL

Python

TCP/IP

Terraform

🕒 August 11

Salve.Inno

11 - 50

💼 Consulting

📣 Marketing

📦 Logistics

Senior SRE operating AWS and Kubernetes cloud platforms for Salve.Inno Consulting’s global customers. Improving observability, automation, reliability, and incident response.

Ansible

AWS

DNS

Grafana

Kubernetes

Linux

MySQL

NoSQL

Postgres

Prometheus

Python

Redis

TCP/IP

Terraform

VoIP

Go

🕒 August 11

Satellite Office

1001 - 5000

📦 Logistics

📣 Marketing

✈️ Travel

DevOps Engineer automating AWS infrastructure and CI/CD pipelines for Satellite Office’s Philippines-based offshore teams. Improving platform security, reliability, scalability, and deployment efficiency.

Ansible

AWS

Azure

Cloud

Docker

EC2

Grafana

Jenkins

Kubernetes

Linux

Prometheus

Python

Terraform

🕒 July 27

TASQ Staffing Solutions

11 - 50

🤝 B2B

🎯 Recruiter

👥 HR Tech

Lead Azure DevOps Engineer implementing cloud solutions in Azure for a client-facing role. Delivering and supporting technical implementations with a focus on client satisfaction.

Ansible

AWS

Azure

Cloud

Google Cloud Platform

Groovy

Jenkins

Linux

Python

Terraform

TFS