Search Remote Jobs

Site Reliability Engineer

Job not on LinkedIn

đŸ”„ 30 minutes ago

đŸ‡ș🇾 United States – Remote

đŸ’” $130k - $150k / year

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🩅 H1B Visa Sponsor

infoinfo

đŸ‘» Ghost score 0%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Accela

Accela

201 - 500 employees

Founded 2000

đŸ’Œ Consulting

📩 Logistics

đŸ—ïž Construction

Consulting ‱ Logistics ‱ Construction

Accela is a leader in providing cloud-based solutions designed to modernize government services. Their unified suite of innovative applications focuses on building more connected communities through enhanced efficiency and secure data management. Accela empowers local and state governments by streamlining processes such as building permits, licensing, cannabis regulation, and environmental health. With a focus on civic solutions, Accela aims to improve public sector operations by eliminating data silos and facilitating better interactions with residents. By leveraging their SaaS platform, governments can enhance service delivery, increase transparency, and reduce operational costs, leading to significant time and cost savings.

📋 Description

‱ Monitor and maintain production cloud environments for availability, performance, scalability, and reliability ‱ Build, configure, and optimize Datadog monitoring, logging, and alerting capabilities ‱ Develop and maintain dashboards for platform health, resource utilization, and service performance ‱ Configure and tune alerts to identify service degradation and operational issues ‱ Support and optimize Azure infrastructure, including AKS, Azure SQL, Azure Storage, and Azure Front Door ‱ Execute production releases, operational changes, and platform maintenance ‱ Participate in and lead incident response activities ‱ Diagnose and resolve production incidents across infrastructure, platform, and application components ‱ Conduct root cause analysis and implement corrective and preventative actions ‱ Maintain operational documentation, runbooks, and incident response procedures ‱ Implement automation to improve operational efficiency and reliability ‱ Provide Level 3 support for customer-reported incidents and service requests ‱ Collaborate with Engineering, Professional Services, and Operations teams on complex technical issues ‱ Meet SLAs for ticket response and resolution ‱ Support customer provisioning, environment maintenance, platform upgrades, and cloud migrations ‱ Support production-to-non-production data refreshes ‱ Assist with database operations, troubleshooting, and data movement ‱ Validate operational activities and maintain documentation and records ‱ Support critical production activities outside normal business hours when required, including customer go-lives and cloud migrations

🎯 Requirements

‱ Bachelor's degree in Computer Science, Computer Engineering, or a related technical field ‱ Minimum 3 years of experience in Site Reliability Engineering (SRE), Cloud Operations, DevOps, Infrastructure Operations, or a related role ‱ Minimum 3 years of hands-on experience supporting production cloud environments ‱ Experience administering and supporting Microsoft Azure cloud services ‱ Hands-on experience with Azure Kubernetes Service (AKS), Azure SQL, Azure Storage, and Azure Front Door ‱ Experience implementing and supporting monitoring, logging, and observability solutions using Datadog ‱ Experience leading incident response activities, performing root cause analysis, and developing operational runbooks ‱ Strong understanding of infrastructure and application performance monitoring, including compute, memory, storage, networking, and availability metrics ‱ Experience supporting SaaS platforms in production environments ‱ Strong troubleshooting, analytical, and problem-solving skills ‱ Excellent written and verbal communication skills with the ability to collaborate effectively across technical and non-technical teams

đŸ–ïž Benefits

‱ Annual bonus target (discretionary, based on company and individual goal achievement) ‱ Flexible time off ‱ Comprehensive medical, dental, and vision plans ‱ Family planning benefits ‱ 401(k) retirement savings plan with company match ‱ Health savings account with company contributions ‱ Flexible spending account ‱ Life, accident, and disability coverage ‱ Business travel insurance ‱ Employee assistance programs ‱ Other well-being benefits ‱ Job accommodations available upon request

Apply Now

Similar Jobs

đŸ”„ 2 hours ago

SGS

10,000+ employees

đŸ’Œ Consulting

đŸ„ Healthcare

đŸœïž Food & Beverage

DevOps Engineer II managing cloud infrastructure, CI/CD, monitoring, and production operations for the U.S. Department of Health and Human Services’ training platform. Supporting secure, reliable application delivery for the Office of Head Start TTA Hub.

đŸ”„ 11 hours ago

Worth AI

11 - 50

đŸ’Œ Consulting

đŸ›Ąïž Insurance

đŸ€– Artificial Intelligence

Senior DevOps Engineer improving Worth AI’s AWS infrastructure, Kubernetes reliability, and CI/CD automation. Building resilient, secure systems that make software delivery faster and easier for engineering teams.

đŸ”„ 12 hours ago

Stratus

501 - 1000

đŸ›Ąïž Insurance

đŸ’Œ Consulting

📩 Logistics

Senior SRE operationalizing reliability for Stratus, whose software digitizes MEP contractor workflows. Building SLOs, observability, incident response, and performance engineering across Azure-based production systems.

đŸ”„ 16 hours ago

Fleetio

51 - 200

📩 Logistics

đŸ’Œ Consulting

🚗 Transport

Senior Site Reliability Engineer scaling Fleetio’s Ruby on Rails fleet-management platform. Improving infrastructure reliability, database performance, observability, and AI-powered operations.

đŸ”„ 17 hours ago

Scribe

51 - 200

☁ SaaS

⚡ Productivity

🏱 Enterprise

Senior DevOps Engineer scaling AWS, Kubernetes, and deployment systems for Scribe’s workflow intelligence platform. Ensuring reliability, observability, and cost-efficient infrastructure.