Senior Site Reliability Engineer – FedRAMP

🔥 0 minutes ago

🇺🇸 United States – Remote

💵 $130k - $160k / year

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

👻 Ghost score 0%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Delinea

Delinea

1001 - 5000 employees

Founded 2021

🔒 Cybersecurity

☁️ SaaS

🏢 Enterprise

💰 Private Equity Round on 2021-03

Cybersecurity • SaaS • Enterprise

Delinea is a cybersecurity company that provides privileged access management (PAM) and identity security software. The Delinea Platform (powered by Iris AI) continuously discovers and inventories human, machine, and AI identities across cloud, on-premise, hybrid, and SaaS environments; vaults, rotates, and governs credentials; enforces just-in-time and zero standing privilege; performs identity posture and threat analysis; and provides real-time, policy-based runtime authorization and auditing for privileged access and AI agents.

📋 Description

• Own reliability for production SaaS services end to end, including availability, performance, and capacity • Define service-level indicators, service-level objectives, and error budgets • Build and tune monitoring in Datadog and Azure Monitor, including detection monitors, synthetic checks, dashboards, and alert routing • Automate incident response and replace manual runbook steps with code • Participate in on-call rotation and lead incident response for high-severity events • Coordinate incident resolution across Support, Engineering, and Product • Write post-incident reviews and customer-facing root cause analyses • Drive preventive actions to completion • Work support escalations by reproducing issues, diagnosing logs, traces, and network captures, and routing or resolving issues with evidence • Build and maintain infrastructure as code with Terraform and Azure DevOps pipelines • Administer the web application firewall, including rule tuning, rate limiting, and false-positive triage • Manage observability platform costs, including Datadog indexing, retention, custom metrics, APM, log ingestion, and archived log storage • Operate within the FedRAMP High environment under change control and continuous monitoring processes • Improve on-call rotation design, escalation paths, alert quality, runbook coverage, and regional handoffs • Partner with Support, Security, Product, and Development to launch services with monitoring, runbooks, and SLOs

🎯 Requirements

• 8+ years in Site Reliability Engineering, DevOps, cloud operations, or production engineering for a SaaS product • Hands-on Azure experience across AKS, App Service, Azure SQL, Redis, Service Bus, Front Door, and Storage, including cloud networking and cloud security fundamentals • Production experience with an observability platform such as Datadog, including metrics, logs, APM, dashboards, and monitor design • Demonstrated ownership of SLIs, SLOs, and error budgets • Incident response experience, including running incident bridges and writing postmortems • Kubernetes production experience, including ingress, deployments, resource limits, and troubleshooting failing workloads • Infrastructure as code with Terraform • CI/CD pipeline creation and troubleshooting; Azure DevOps preferred • Scripting in PowerShell and Python • Fluency with YAML and JSON • Strong networking and web fundamentals: DNS, TLS and certificate chains, load balancing, reverse proxies, firewalls, and packet-level troubleshooting • Knowledge of redundancy, backup, and disaster recovery strategies in cloud environments • Clear written communication • Willingness to participate in an on-call rotation covering weekends and emergencies • Up to 10% travel • Prior government cloud experience is welcome but not required • Must not require any type of U.S. work authorization now or in the future

🏖️ Benefits

• Equity • Performance-based bonus program or role-based incentive programs • Healthcare insurance • Pension/retirement matching • Comprehensive life insurance • Employee assistance program • Time off plans • Paid company holidays • Career progression • Meaningful work and culture of innovation

Apply Now

Similar Jobs

🔥 1 hour ago

LexisNexis

10,000+ employees

⚖️ Legal

💼 Consulting

🏥 Healthcare

Senior SRE modernizing Azure platforms, observability, and disaster recovery for LexisNexis patent data systems. Driving reliability, incident response, automation, and AI-assisted operations.

🔥 1 hour ago

RELX

10,000+ employees

💼 Consulting

🏥 Healthcare

🛡️ Insurance

Senior SRE modernizing Azure platforms and incident response for LexisNexis patent data solutions. Driving reliability, disaster recovery, automation, and AI-assisted operations.

🔥 5 hours ago

Avanade

10,000+ employees

💼 Consulting

📦 Logistics

📣 Marketing

Avanade manager architecting Azure DevOps, GitHub, and AI-enabled software delivery solutions for enterprise clients. Leading DevOps transformation, Copilot adoption, governance, and cloud engineering modernization.

🔥 17 hours ago

Akkadian Labs

51 - 200

☁️ SaaS

🏢 Enterprise

📡 Telecommunications

Senior DevOps Engineer scaling secure AWS, hybrid-cloud, and AI infrastructure for Akkadian Labs’ enterprise collaboration automation platform. Improving deployments, observability, reliability, and operational governance.

🔥 19 hours ago

Casper Studios

2 - 10

🤖 Artificial Intelligence

🏢 Enterprise

📱 Media

AI Security & DevOps Engineer securing AI systems from prototype to production for Casper Studios, an AI services firm. Owning security standards, platform controls, and enterprise client approvals.