DevOps Reliability Engineer

🕒 June 10

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Advanced Solutions International, Inc.

Advanced Solutions International, Inc.

201 - 500 employees

Founded 1991

☁️ SaaS

🤝 B2B

📚 Education

🔥 Funding within the last year

💰 Private Equity Round - Advanced Solutions International on 2025-10

SaaS • B2B • Education

Advanced Solutions International, Inc. is a software company that provides a suite of solutions for associations and non-profit organizations, centered on its iMIS engagement management system. ASI offers products including iMIS, TopClass Learning Management, OpenWater application & review, Clowder mobile engagement, SpaceMaster advertising management, and other integration and data-management tools; it markets these as a combined platform (iMIS Power Suite) to help non-profits manage members, learning, events, and data. The company runs partner programs, demo events, and conferences aimed at the association/non-profit market and emphasizes SaaS delivery, integrations, and client support.

📋 Description

• Monitor and improve the health, availability, performance, and cost efficiency of Azure-based production systems. • Use application, database, and infrastructure telemetry to identify performance issues, bottlenecks, and reliability risks. • Tune Azure services and platform configurations to maximize performance, resilience, and resource efficiency. • Partner with engineering teams to recommend and implement practical, data-driven improvements to reliability, scalability, and operational effectiveness. • Create and maintain operational documentation, runbooks, and troubleshooting guides to support consistent incident response and ongoing operations. • Support Tech Support and Sustained Engineering by executing approved SQL queries and completing database backups and restores for troubleshooting purposes. • Analyze how partner integrations and customer usage patterns impact system performance and cloud spend. • Investigate complex production issues, perform root cause analysis, and drive resolution of reliability and performance problems. • Contribute to continuous improvement across deployment processes, system stability, and operational readiness. • Perform other job-related duties and responsibilities as assigned.

🎯 Requirements

• Bachelors degree in Computer Science, Information Technology or related degree or relevant experience. • 8+ years of experience in DevOps, Site Reliability Engineering, Cloud Engineering, or similar roles. • Strong hands-on experience with Microsoft Azure, especially: Azure SQL, Azure Functions, Azure App Services, and Azure Containers (AKS, Container Apps, or similar). • Ability to read and interpret telemetry, logs, metrics, and resource usage data and explain what’s wrong and how to fix it. • Experience working with production systems that require high availability and reliability. • Comfort owning work end-to-end, from identifying issues to executing improvements. • Experience adjusting pipelines, hosting configurations, and deployment processes. • Solid understanding of cloud cost drivers and usage optimization. • Strong problem-solving skills and the ability to work collaboratively across engineering and support team. • Ability to read and interpret application code to support troubleshooting, root cause analysis, and identification of performance improvement opportunities.

🏖️ Benefits

• Wellness Benefits • Opportunities for Professional Growth and Development • Flexible Remote Work • Volunteer Time Off • Study Leave • Employee Assistance Program

Apply Now

Similar Jobs

🕒 June 4

Omilia - Conversational Intelligence

201 - 500

💼 Consulting

🛡️ Insurance

✈️ Travel

Senior Site Reliability Engineer maintaining production clusters and developing observability solutions. Collaborate with teams to ensure platform reliability and performance using automation and monitoring tools.

Ansible

AWS

Cloud

Docker

Grafana

Kubernetes

Linux

MySQL

NoSQL

Postgres

Prometheus

Python

RDBMS

Redis

TCP/IP

Terraform

VoIP

Go

🕒 April 2

ClickHouse

51 - 200

☁️ SaaS

🏢 Enterprise

🤖 Artificial Intelligence

Database Reliability Engineer driving improvements in performance and reliability for ClickHouse. Collaborating with global teams to optimize operations and enhance service reliability.

AWS

Azure

Cloud

Google Cloud Platform

Python

SQL

🕒 March 28

RevenueCat

51 - 200

💼 Consulting

📣 Marketing

☁️ SaaS

Senior DevOps/DevEx Engineer responsible for building internal development tools at RevenueCat. Collaborating with a global remote team across diverse geographic locations.

AWS

Cloud

Docker

Kubernetes

Python

🕒 March 16

Binance

1001 - 5000

₿ Crypto

💳 Fintech

Senior DevOps Engineer or Architect elevating infrastructure management in Binance's blockchain ecosystem. Architecting solutions and improving deployment processes in a collaborative environment.

Ansible

AWS

Cloud

Docker

ElasticSearch

Google Cloud Platform

Kafka

Linux

Python

Terraform

Go

🕒 March 13

ClickHouse

51 - 200

☁️ SaaS

🏢 Enterprise

🤖 Artificial Intelligence

Senior Site Reliability Engineer at ClickHouse leading reliability initiatives for cloud infrastructure. Collaborating with engineering teams to design and implement scalable, fault-tolerant systems.

Ansible

AWS

Azure

Cloud

Docker

Google Cloud Platform

Kubernetes

Puppet

Python

SQL

Terraform

Go