Senior Cloud Ops Engineer – SRE

🕒 July 17

🇺🇸 United States – Remote

💵 $90k - $125k / year

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of CaseWorthy, Inc.

CaseWorthy, Inc.

51 - 200 employees

🏥 Healthcare

💼 Consulting

📦 Logistics

Healthcare • Consulting • Logistics

CaseWorthy, Inc. is a mission-critical case management platform designed to deliver comprehensive care and outcome reporting for social and human services. It provides centralized data management through its CORE platform, enabling organizations to streamline operations, enhance service delivery, and gain 360-degree views of client and program data. CaseWorthy supports a wide range of programs and empowers organizations to transform communities, offering solutions tailored to needs in aging services, behavioral health, veterans services, and more. With a strong focus on data-driven insights and operational efficiency, CaseWorthy aims to empower health and human service organizations to achieve their missions through innovative technology.

📋 Description

• Maintain and improve the company's Azure cloud infrastructure with a focus on reliability, availability, and performance. • Define, track, and report on SLIs, SLOs, and error budgets for production services. • Provide on-call, after-hours support on a rotating schedule, including incident response and escalation. • Lead and participate in blameless postmortems and retrospectives of production infrastructure incidents, driving corrective and preventive action items to closure. • Perform hosting tasks as needed including deployments and patching. • Automate infrastructure and configuration management using Infrastructure as Code (IaC). • Maintain and execute new tenant site provisioning. • Plan and execute database migrations to accommodate capacity plans. • Build and maintain observability tooling — metrics, logs, traces, dashboards, and alerting — using Datadog and Azure Monitor to enhance security, observability, and monitoring of workloads (compute, data storage, application integration elements). • Perform capacity planning and load/performance analysis to anticipate and prevent reliability issues before they impact customers. • Contribute to the organization's security audits and risk assessments. • Assist with vulnerability scans / penetration tests for internal and client systems. • Assist with identifying, documenting, and socializing application risks and vulnerabilities. • Use Azure services such as Azure Virtual Machines, Azure Kubernetes Service (AKS), Azure Functions, Azure SQL Database / Azure SQL Managed Instance, Azure Cosmos DB, Azure Blob Storage, Azure API Management, and Azure DNS, to name a few. • As well as Azure DevOps Pipelines / GitHub Actions, Azure Resource Manager (ARM) templates / Bicep, Checkmarx, Tenable, Datadog, Azure Monitor, and Atlassian products. • Ability to travel nationwide, up to 10% annually. • Perform other duties as assigned.

🎯 Requirements

• 3+ years of experience working with public cloud infrastructure, specifically Microsoft Azure. • Background in site reliability engineering (SRE), DevSecOps, or software development. • Hands-on experience with observability and monitoring platforms, specifically Datadog and Azure Monitor. • Working knowledge of SRE fundamentals — SLIs, SLOs, error budgets, incident response, and blameless postmortems. • Experience with deployment pipeline automation tools. • Experience with scripting languages (Python, PowerShell). • Basic knowledge of SQL. • B.S. in IT, Computer Science, or related field. • Experience with Azure-hosted environments. • Familiarity with Windows systems administration. • Experience with container orchestration (Azure Kubernetes Service / Kubernetes). • Microsoft Azure certifications (e.g., AZ-104, AZ-400, AZ-500) are a plus. • Datadog certification is a plus. • Experience with AWS or GCP is a plus, given transferable public cloud skills. • M.S. in IT, Computer Science, or related field.

Apply Now

Similar Jobs

🕒 July 17

Siemens Healthineers

10,000+ employees

🏥 Healthcare

⚕️ Healthcare Insurance

🧬 Biotechnology

Network Engineer responsible for cloud solutions operations at Varian. Collaborating across teams to enhance, optimize, and maintain cloud computing capabilities.

🇺🇸 United States – Remote

💰 $1.5M Grant on 2021-05

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

Cloud

Firewalls

Switching

TCP/IP

🕒 July 17

Siemens Healthineers

10,000+ employees

🏥 Healthcare

⚕️ Healthcare Insurance

🧬 Biotechnology

Network Engineer managing cloud operations and support for Varian’s cloud solutions. Enhancing, optimizing, and maintaining computing capabilities across the global cloud solution.

🇺🇸 United States – Remote

💰 $1.5M Grant on 2021-05

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

Azure

Citrix

Cloud

Firewalls

Switching

TCP/IP

🕒 July 16

Miris

11 - 50

☁️ SaaS

🥽 AR/VR

🤝 B2B

Site Reliability Engineer building scalable platforms for 3D/4D content delivery at Miris. Collaborating with teams to ensure system reliability and performance across AR/VR devices.

🇺🇸 United States – Remote

💵 $102.7k - $287.5k / year

💰 $26M Seed Round - MIRIS on 2024-08

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🕒 July 16

Epic Kids

11 - 50

📚 Education

👥 B2C

📱 Media

Senior Site Reliability Engineer driving reliability and stability for Epic Kids' GCP infrastructure. Collaborating with product and data teams to ensure platform efficiency and security.

🇺🇸 United States – Remote

💵 $160k - $200k / year

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🕒 July 16

Global Alliant Inc

51 - 200

💼 Consulting

🏥 Healthcare

📦 Logistics

Fullstack DevSecOps Engineer developing enterprise applications on AWS. Collaborating in Agile environments with strong focus on secure coding practices.