Search Remote Jobs

Lead Site Reliability Engineer

đź•’ June 17

🇺🇸 United States – Remote

đź’µ $160k - $185k / year

⏰ Full Time

đźź  Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🦅 H1B Visa Sponsor

info
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Accela

Accela

201 - 500 employees

Founded 2000

đź’Ľ Consulting

📦 Logistics

🏗️ Construction

Consulting • Logistics • Construction

Accela is a leader in providing cloud-based solutions designed to modernize government services. Their unified suite of innovative applications focuses on building more connected communities through enhanced efficiency and secure data management. Accela empowers local and state governments by streamlining processes such as building permits, licensing, cannabis regulation, and environmental health. With a focus on civic solutions, Accela aims to improve public sector operations by eliminating data silos and facilitating better interactions with residents. By leveraging their SaaS platform, governments can enhance service delivery, increase transparency, and reduce operational costs, leading to significant time and cost savings.

đź“‹ Description

• Serve as a technical leader for reliability engineering, operational excellence, and platform modernization across the Civic Platform. • Drive platform modernization initiatives, including the continued evolution from VM-based architectures toward containerized and cloud-native services, in partnership with DevOps Engineering, Database Engineering, Security, and Development teams. • Lead efforts that improve and sustain the availability, performance, scalability, security, and cost efficiency of Accela's SaaS offerings. • Define, implement, and operate service level objectives (SLOs), service level agreements (SLAs), and error budgets for critical platform services, using data to drive prioritization and risk-based decision making. • Lead observability initiatives across metrics, distributed tracing, logging, and monitoring platforms to improve system visibility and accelerate issue detection and resolution. • Drive Root Cause Analysis (RCA) efforts for complex production incidents, facilitate blameless postmortems, and ensure corrective actions are implemented and tracked to completion. • Design, develop, and maintain automation, tooling, and software solutions that improve reliability, operational efficiency, scalability, and developer productivity. • Serve as a senior technical escalation point during production incidents and for platform changes that impact availability, performance, security, or compliance. • Partner with Security and Compliance teams to ensure platform operations meet regulatory and compliance requirements, including SOC 2, HIPAA, FedRAMP, StateRAMP, and PCI-DSS. • Translate operational metrics, reliability trends, and platform health data into actionable insights for engineering leadership and executive stakeholders. • Mentor engineers across the Cloud Engineering organization and influence engineering best practices through technical leadership and collaboration

🎯 Requirements

• 8+ years of experience in Site Reliability Engineering, Software Engineering, Cloud Infrastructure, or related disciplines within a SaaS environment, including experience leading complex technical initiatives. • Demonstrated technical leadership driving platform modernization in containerized and orchestrated environments, including Kubernetes or equivalent technologies. • Hands-on experience operating and supporting large-scale SaaS platforms on Microsoft Azure. • Experience developing automation and operational tooling using Python, PowerShell, Bash, or similar scripting languages. • Deep expertise designing, operating, analyzing, and troubleshooting complex distributed systems across the application, infrastructure, networking, and operating system layers. • Strong experience with modern observability platforms, including monitoring, logging, metrics, and distributed tracing. • Demonstrated success leading incident response, Root Cause Analysis, and continuous improvement initiatives. • Experience establishing and maturing Incident, Problem, and Change Management practices. • Strong written and verbal communication skills with the ability to effectively communicate technical concepts to engineering leadership and executive stakeholders. • Experience using Git and GitHub-based development workflows.

🏖️ Benefits

• flexible time off • comprehensive medical, dental, and vision plans • family planning benefits • 401(k) retirement savings plan with company match • health savings account with company contributions • flexible spending account • life, accident, and disability coverage • business travel insurance • employee assistance programs • other well-being benefits

Apply Now

Similar Jobs

đź•’ June 17

Clinician Nexus

51 - 200

🏥 Healthcare

⚕️ Healthcare Insurance

📚 Education

Senior Manager, DevOps enabling health organizations with technology while leading a team of engineers. Overseeing CI/CD, infrastructure automation, and collaborating with multiple stakeholders.

đź•’ June 17

ClassWallet

11 - 50

đź’ł Fintech

📚 Education

🏛️ Government

DevOps Engineer optimizing cloud infrastructure and deployment pipelines for fintech company. Redefining public funds management and ensuring system reliability with high compliance standards.

đź•’ June 16

Virtusa

10,000+ employees

đź’Ľ Consulting

🏢 Enterprise

🤖 Artificial Intelligence

Lead DevOps Engineer at Virtusa designing high-availability infrastructure with OpenShift and Proxmox technologies. Focusing on security standards, automation, and storage integration.

🇺🇸 United States – Remote

💵 zł26.4k - zł31k / year

đź’° $108M Post-IPO Equity - Virtusa on 2017-05

⏰ Full Time

đźź  Senior

⛑ DevOps & Site Reliability Engineer (SRE)

đź•’ June 16

SailPoint

1001 - 5000

đź’Ľ Consulting

🏥 Healthcare

📦 Logistics

CI/CD Engineering Manager responsible for nurturing a high-performing team in a cloud-native SaaS environment. Leading security, compliance, and operational excellence for developer enablement and platform engineering.

đź•’ June 16

DexCare

51 - 200

🏥 Healthcare

đź’Ľ Consulting

📦 Logistics

Senior Site Reliability Engineer at DexCare managing cloud-native infrastructure in AWS. Designing secure systems and ensuring observability while collaborating in an Agile environment.