Manager, Site Reliability Engineering

Job not on LinkedIn

🔥 18 minutes ago

🇺🇸 United States – Remote

💵 $160k - $180k / year

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

👻 Ghost score 17%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Climb Channel Solutions NA

Climb Channel Solutions NA

51 - 200 employees

Founded 1982

💼 Consulting

📦 Logistics

📣 Marketing

Consulting • Logistics • Marketing

Climb Channel Solutions NA is an IT distribution company focused on providing leading and innovative technology solutions. They support technology resellers by offering expertise in areas such as virtualization, cloud, data management, and cybersecurity, thereby enhancing their partners' success. Climb is dedicated to transforming IT distribution with exceptional service and an extensive vendor marketplace to facilitate business growth for its partners across various sectors, including public and private markets.

📋 Description

• Lead the Platform SRE and DevOps team supporting Delinea products • Review pull requests and validate pipeline changes • Tune monitors and dashboards, run Datadog queries, and troubleshoot production issues • Own availability and performance of production environments across Azure and AWS • Manage AKS workloads, ingress and networking, data services, messaging, and CDN/WAF layers • Hire, onboard, coach, and develop full-time SRE engineers • Direct and manage contractor resources, including scoping work and reviewing deliverables • Lead one-on-ones, standups, and planning sessions across time zones • Participate in the on-call rotation and serve as incident commander for Sev1 and Sev2 events • Own detection, triage, mitigation, customer-facing status communication, and post-incident reviews • Ensure RCAs meet customer-ready standards and drive preventative actions to closure • Improve observability through SLI/SLO definition, alert-quality improvements, synthetic coverage, APM instrumentation, log hygiene, and dashboard standards • Grow into supporting Delinea’s FedRAMP High environment • Reduce operational toil through automation • Report incident metrics, trends, and reliability commitments to leadership and create improvement plans

🎯 Requirements

• 6+ years in Site Reliability Engineering, DevOps, or Cloud Operations, with demonstrated ownership of production SaaS systems • 2+ years of direct people leadership, including performance management, hiring, and coaching • Current, hands-on production experience with Azure Kubernetes Service, core Azure services (SQL, Redis, Service Bus, Blob Storage), AWS services (SES, EC2, RDS), WAF, Azure DevOps pipelines, Datadog, and Atlassian Jira Service Management • Hands-on experience across both Azure and AWS is required • Deep observability expertise covering metrics, logs, traces, synthetics, SLOs, and alerting strategy • Hands-on proficiency with Datadog or an equivalent platform, including APM trace analysis and log-based troubleshooting • Proven experience serving as incident commander for major incidents • Strong cloud networking and security fundamentals: load balancing, DNS, TLS and certificate lifecycle, firewalls, VPN, routing, and identity and access management • Automation and scripting ability in PowerShell, Python, Bash, or similar • Practical infrastructure-as-code experience with Terraform, ARM, or Bicep • Practical experience with multi-region, multi-tenant SaaS architectures, including backup, redundancy, and disaster recovery • Excellent written communication for customer-facing status updates and incident summaries • Willingness and availability to work across time zones and participate in an on-call rotation • U.S. work authorization required; Delinea will not consider candidates needing current or future U.S. work authorization sponsorship • Direct experience operating in a FedRAMP or other regulated environment is preferred • Experience with incident management programs, public status pages, customer notifications, Atlassian Jira Service Management, Confluence, Azure DevOps, and cloud cost optimization is preferred

🏖️ Benefits

• Equity • Performance-based bonus or role-based incentive programs • Healthcare insurance • Pension/retirement matching • Comprehensive life insurance • Employee assistance program • Time off plans • Paid company holidays • Meaningful work • Career progression • Global team environment

Apply Now

Similar Jobs

🔥 9 hours ago

Rockstar

1 - 10

💼 Consulting

📣 Marketing

📦 Logistics

DevSecOps Engineer securing and scaling AWS, Kubernetes, and infrastructure-as-code environments. Supporting compliance, production reliability, and security-sensitive engineering platforms.

🔥 13 hours ago

Bloomerang

201 - 500

💼 Consulting

📣 Marketing

🤝 Non-profit

Senior SRE improving reliability, observability, and automation for Bloomerang’s nonprofit fundraising platform. Troubleshooting production systems and leading incident response across application and infrastructure stacks.

🔥 14 hours ago

High 5 Games

51 - 200

🎮 Gaming

🎲 Gambling

🤝 B2B

DevOps Engineer scaling Google Cloud infrastructure for machine learning operations. Automating deployments, pipelines, monitoring, and reliability for AI systems serving millions of players.

🔥 20 hours ago

CoorsTek, Inc.

5001 - 10000

🏭 Manufacturing

🚘 Automotive

🎖️ Defense

AI Deployment Engineer delivering secure production applications and AI automation for CoorsTek’s advanced ceramics manufacturing operations. Translating plant and business needs into scalable software.

🔥 20 hours ago

Particle41

51 - 200

💼 Consulting

📣 Marketing

☁️ SaaS

Lead DevOps Engineer owning cloud infrastructure, automation, reliability, and security. Advising Particle41 clients and mentoring teams delivering complex digital projects.