Senior Site Reliability Engineer

Job not on LinkedIn

🔥 3 minutes ago

🇺🇸 United States – Remote

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

👻 Ghost score 12%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of NeuroFlow

NeuroFlow

51 - 200 employees

Founded 2017

🏥 Healthcare

☁️ SaaS

🤝 B2B

💰 $2.8M Grant - NeuroFlow on 2025-05

Healthcare • SaaS • B2B

NeuroFlow is a healthcare technology company that provides a cloud-based behavioral health infrastructure and analytics platform for health plans, health systems (ACOs & IDNs), federal agencies, and behavioral health providers. The platform captures assessments and survey data at scale, uses transparent machine learning on claims, EHR, ADT and assessment data to identify and stratify behavioral health risk, engages members with acuity-driven next best actions, and manages referrals and outcomes tracking to improve access, reduce acute utilization, and demonstrate ROI. The company emphasizes measurement-informed care, HIPAA/HITRUST/SOC2 compliance, and has published peer-reviewed evidence showing reductions in total medical PMPM, inpatient admissions, and ED visits, as well as improved referral and assessment metrics.

📋 Description

• Own the reliability of production systems and how reliability is measured • Define SLOs and SLIs with stakeholders, establish error budgets, and use them to guide release decisions • Run and improve the SaaS monitoring, alerting, and logging stack to meet compliance requirements and detect issues proactively • Participate in on-call production incident response and help engineers resolve customer issues • Lead blameless postmortems, identify root causes, and implement lasting fixes • Identify, measure, and automate manual repetitive engineering work • Maintain automation and documentation that survives ownership changes • Own the design, build, and maintenance of core infrastructure for safe and efficient product delivery • Set priorities for reliability and platform work, obtain manager approval, and adapt to business needs • Assess risk versus impact for high-visibility systems and roll out measurable, reversible improvements • Make build-versus-buy decisions and recommend proven tools • Partner with engineering teams on CI/CD, testing, canary releases, and automated rollback practices • Work with security to maintain infrastructure and operations compliance and proactively raise risks • Mentor engineers on reliability, Azure, AWS, Terraform, and operational practices • Create documentation for system operation and lead solutions with clear tradeoffs

🎯 Requirements

• 8+ years in site reliability, DevOps, or infrastructure engineering, including ownership of production systems from end to end • Built and managed AWS and Azure accounts and resources in line with the Well-Architected Framework • Experience with AWS services including ECS, Fargate, RDS, Lambda, SNS, SQS, S3, EventBridge, and Step Functions • Deployed Docker-based software to production and understanding of reliable container operations • Hands-on experience across Azure, AWS, and traditional data centers, including managing Windows VMs and IIS • Deep expertise in SQL and relational database administration, including query tuning, index optimization, and resolving high-load production incidents; SQL Server preferred • Infrastructure-as-code experience using Terraform • Practical experience with DevOps principles, the 12-Factor App, least privilege access, and zero-trust architecture • Experience leading cross-team requirements and SLA definition • Experience defining and operating against SLOs and leading incident response and postmortems for production outages • Experience with ITIL-aligned service management, including incident, problem, and change management • Experience mentoring engineers on reliability and delivery practices • United States Citizenship • Ability to satisfy security investigation and eligibility requirements for access to classified (Public Trust) information • Azure or AWS certifications preferred, such as AZ-104, AZ-305, or AWS Certified Solutions Architect • Experience with Azure Government or other federal cloud environments preferred • Experience supporting the VA, DoD, or other federal health programs, including FedRAMP or ATO work preferred • PowerShell scripting and automation preferred • Experience in a HIPAA-compliant environment preferred • Security experience beyond day-to-day operations, such as red teaming or penetration testing, preferred

🏖️ Benefits

• Flexible work schedule • Unlimited PTO • Physical and mental wellness benefits • Medical coverage • Parental leave • 401K • Company-sponsored events • Referral program • Onsite gym • Dog friendly office • Snacks in the office • Commuter benefits • Onsite massages

Apply Now

Similar Jobs

🔥 1 hour ago

RTX

10,000+ employees

🏭 Manufacturing

💼 Consulting

📦 Logistics

Reliability Engineer improving equipment reliability and maintenance strategies across RTX aerospace manufacturing sites. Applying RCA, FMEA, CMMS/EAM analytics and predictive-maintenance practices.

🔥 2 hours ago

Ad Hoc LLC

501 - 1000

💼 Consulting

🏥 Healthcare

📦 Logistics

DevOps Engineer III building AWS and Kubernetes infrastructure for Ad Hoc’s government digital services. Improving CI/CD, security, reliability, and developer experience for Veterans Affairs products.

🔥 3 hours ago

Circle

501 - 1000

💳 Fintech

₿ Crypto

🌐 Web 3

Senior SRE operating secure Kubernetes and Terraform infrastructure for Circle’s digital-asset financial platform. Improving reliability, observability, automation, and production resilience.

🔥 7 hours ago

RELX

10,000+ employees

💼 Consulting

🏥 Healthcare

🛡️ Insurance

SRE Engineering Lead managing teams and reliability initiatives for LexisNexis Risk Solutions' cloud-based risk platforms. Driving Kubernetes, Terraform, Azure, automation, incident response, and operational excellence.

🔥 7 hours ago

RELX

10,000+ employees

💼 Consulting

🏥 Healthcare

🛡️ Insurance

Site Reliability Engineering Lead managing SRE teams and cloud reliability for LexisNexis Risk Solutions. Driving Kubernetes, Terraform, Azure, automation, incident response, and resilient platforms.