Site Reliability Engineer

Job not on LinkedIn

🔥 3 minutes ago

🇨🇦 Canada – Remote

💵 $80k - $110k / year

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

👻 Ghost score 0%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Rentsync

Rentsync

51 - 200 employees

Founded 2010

💼 Consulting

📣 Marketing

📦 Logistics

💰 Private equity on 2025-06

Consulting • Marketing • Logistics

Rentsync is a North American proptech SaaS company that provides an integrated platform for multifamily property marketing, leasing, and resident management. Its products include marketing automation and AI tools, listing syndication, customizable websites, digital applications and payments, tenant portals, analytics, and agency digital marketing services to help property managers attract, qualify, convert, and coordinate tenants across the full lease lifecycle. Rentsync primarily serves property management and multifamily marketing teams, offering APIs and integrations with top industry partners to streamline workflows and improve performance.

📋 Description

• Act as first responder for production alerts and incidents across services, from triage through resolution • Diagnose and fix issues directly in AWS and Kubernetes, including failing pods, resource exhaustion, bad deployments, networking/DNS, database, and cache problems • Roll back, scale, reconfigure, or patch infrastructure to restore service quickly • Escalate to development teams only when a code change is needed, providing a clear diagnosis • Own PagerDuty setup and business-hours incident response while reducing MTTD and MTTR • Run blameless post-mortems and drive technical follow-up work • Automate runbooks and repetitive operational work, including AI-assisted triage, investigation, and remediation • Build and maintain monitoring for Kubernetes workloads and services using Prometheus/Mimir, Loki, Tempo, Grafana, and OpenTelemetry • Monitor production releases and identify regressions in latency, errors, or resource usage • Create and maintain synthetic checks, smoke tests, health checks, and load/performance tests • Partner with engineering teams on performance and reliability issues and define SLOs, SLIs, and error budgets • Harden the platform through Terraform changes, Kubernetes resource tuning, autoscaling, CI/CD checks, secrets, and IAM improvements • Maintain service documentation and architecture decisions

🎯 Requirements

• 3+ years in a cloud engineering, DevOps, or SRE role supporting production web applications • Hands-on production incident response experience, including diagnosing and fixing issues directly • Strong hands-on AWS production experience with EKS, EC2, RDS, VPC networking, IAM, and CloudWatch • Deep production Kubernetes experience, including troubleshooting, debugging, and monitoring workloads • Experience identifying and resolving performance and reliability issues with engineering teams • Experience with monitoring and observability tools such as Prometheus, Grafana, Loki, Datadog, or CloudWatch • Experience with on-call and alerting tools such as PagerDuty • Experience building automated production reliability tests or checks, including synthetic, smoke, health, or load tests • Willingness to participate in a future after-hours on-call rotation • Infrastructure as code with Terraform and comfort working in CI/CD pipelines • Solid Linux, networking, and container fundamentals • Scripting/automation in Bash, Python, or similar • Calm, clear communication during incidents and across teams • Preferred: AI tools for SRE work; Azure and possibly GCP; multiple technology stacks; AWS certification; LGTM or OpenTelemetry at scale; k6, Locust, or JMeter; MySQL/PostgreSQL operations; Redis/Memcached tuning; Cloudflare; cloud cost optimization and capacity planning

🏖️ Benefits

• Remote position • Preferential consideration may be given to individuals within a reasonable commuting distance of one of the offices • Equal opportunity employer • Unique accommodations available during the interview process • Potential criminal background check in the final interview phase

Apply Now

Similar Jobs

🔥 4 hours ago

commonsku

51 - 200

☁️ SaaS

🤝 B2B

DevOps Engineer II managing AWS infrastructure, IaC, CI/CD, and observability. Supporting commonsku’s SaaS platform for promotional-products distributors.

🕒 4 days ago

Clario

5001 - 10000

🏥 Healthcare

💼 Consulting

📦 Logistics

Senior DevOps Engineer advancing AWS cloud infrastructure, CI/CD, and DevSecOps for Clario's clinical research technology. Driving secure, scalable software delivery across global engineering teams.

🕒 4 days ago

Practice EHR

51 - 200

🏥 Healthcare

💼 Consulting

⚕️ Healthcare Insurance

Intermediate DevOps Engineer building reliable AWS infrastructure and CI/CD automation. Supporting Practice Better’s EHR and practice management platform for health and wellness practitioners.

🕒 September 15

Case IQ

51 - 200

💼 Consulting

🏥 Healthcare

⚖️ Legal

DevOps Engineer optimizing Azure and AWS infrastructure for Case IQ’s governance, risk, and compliance software. Automating deployments, improving reliability, security, observability, and AI-enabled operations.

🕒 September 15

GBL

11 - 50

🎮 Gaming

💳 Fintech

DevOps and Golang Developer building Golang microservices and automated infrastructure. Scaling secure systems for a crypto-first international online casino.