Senior Site Reliability Engineer

🕒 August 19

🇨🇦 Canada – Remote

💵 CA$131k - CA$164.3k / year

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

👻 Ghost score 7%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Blackpoint Cyber

Blackpoint Cyber

51 - 200 employees

💼 Consulting

🎖️ Defense

🔒 Cybersecurity

💰 $190M Series C on 2023-06

Consulting • Defense • Cybersecurity

Blackpoint Cyber is a technology-focused cybersecurity company headquartered in Maryland, USA. Established by former US Department of Defense and Intelligence security experts, Blackpoint leverages its real-world cyber experience to help Managed Service Providers (MSPs) safeguard their infrastructure and operations. The company offers a proprietary cybersecurity ecosystem, including its SNAP-Defense platform for Managed Detection and Response (MDR) services. Blackpoint's dedicated security analysts work 24/7 to combine various security measures, including network visualization and endpoint security, to monitor and respond to threats. Additionally, Blackpoint is launching LogIC, a logging and integrated compliance service designed to assist MSPs with cyber compliance requirements. The company's mission is to deliver comprehensive detection and response services to help MSPs combat the evolving threat landscape.

📋 Description

• Design, develop, and maintain highly scalable infrastructure using Infrastructure as Code (Terraform and Terragrunt) for automated cloud resource provisioning and orchestration • Own and optimize the AWS cloud environment for cost efficiency, security best practices, and high availability • Manage and optimize Kubernetes cluster environments using Helm, ArgoCD, Istio, and Kustomize • Administer and scale data streaming infrastructure using Confluent Cloud and Apache Kafka • Deploy, configure, and maintain Redis for caching and real-time data processing • Implement and maintain monitoring, alerting, and incident response frameworks using Prometheus, Grafana, Alert Manager, and OpsGenie/PagerDuty • Facilitate controlled feature deployments and progressive rollouts through LaunchDarkly/PostHog • Partner with software development teams to integrate new services, applications, and features into existing infrastructure • Diagnose and resolve complex system-level issues while maintaining performance and maximizing uptime • Drive continuous improvement of automation tooling, operational processes, and engineering methodologies • Stay current on emerging SRE trends and tools and help adopt relevant industry advancements and best practices

🎯 Requirements

• 5+ years of experience in a Senior Site Reliability Engineer role or equivalent, with substantial emphasis on cloud infrastructure management and automation • Expertise in Infrastructure as Code using Terraform and Terragrunt for enterprise-scale deployments • Comprehensive knowledge of AWS, including designing, implementing, and maintaining secure, scalable, resilient cloud architectures • Extensive hands-on experience with distributed data streaming using Confluent Cloud and Apache Kafka • Proven experience with Redis for caching and Amazon RDS for relational database management • Experience with enterprise search and analytics platforms including OpenSearch, Elasticsearch, and ChaosSearch • Proficiency designing and implementing monitoring and alerting infrastructure using Prometheus, Grafana, Alert Manager, and OpsGenie/PagerDuty • Practical experience with feature flag systems including LaunchDarkly/PostHog for controlled release management • Extensive experience administering production-grade Kubernetes with Helm, ArgoCD, and Istio; working knowledge of Kustomize • Strong problem-solving skills, with the ability to troubleshoot complex systems in production • Strong communication and collaboration skills, with experience working in Agile environments • Are you authorized to work in Canada without restriction? • Will you now or in the future require sponsorship to work within Canada?

🏖️ Benefits

• Equity participation available to employees globally, with program details varying by location and employment structure • Competitive Health, Vision, Dental, and Life Insurance plans for eligible employees in the US • Robust 401k plan for eligible employees in the US • Discretionary Time Off for eligible employees in the US • Other minor perks • International employees receive competitive benefits in accordance with local market standards and applicable country requirements

Apply Now

Similar Jobs

🕒 August 11

Autodesk

10,000+ employees

🏗️ Construction

🏭 Manufacturing

💼 Consulting

Senior DevOps Developer building reliable AWS, Kubernetes, and MongoDB services for Autodesk Construction Solutions. Improving automation, observability, security, disaster recovery, and production reliability for construction software customers.

🕒 July 29

EXL

10,000+ employees

🏥 Healthcare

🛡️ Insurance

📦 Logistics

Release Engineer managing data and BI releases for EXL’s migration and modernization programs. Coordinating deployments, documentation, support readiness, and post-release resolution.

🇨🇦 Canada – Remote

💵 C$120k - C$127k / year

💰 $2M Venture Round on 2015-01

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🕒 July 28

MaintainX

501 - 1000

☁️ SaaS

🏭 Manufacturing

🏢 Enterprise

Site Reliability Engineer at MaintainX enhancing reliability and observability while collaborating with product teams. Mentoring developers and establishing best practices for system design and operational readiness.

🕒 July 28

Thumbtack

1001 - 5000

🏪 Marketplace

☁️ SaaS

Senior Software Engineer designing resilient systems for availability and scalability at Thumbtack. Collaborating with teams and managing infrastructure effectively.

🇨🇦 Canada – Remote

💵 $180.2k - $233.2k / year

💰 $75M Debt Financing - Thumbtack on 2024-07

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🕒 July 23

Kong Inc.

201 - 500

💼 Consulting

📦 Logistics

🔌 API

Senior Site Reliability Engineer focused on orchestrating cloud-native systems for Kong's Managed Gateways. Leading a team to ensure robust and scalable infrastructure for enterprise customers.

🇨🇦 Canada – Remote

💵 CA$118k - CA$167k / year

💰 $100M Series D on 2021-02

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)