Senior Site Reliability Engineer

🔥 14 hours ago

🇨🇦 Canada – Remote

💵 $145k - $193k / year

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

👻 Ghost score 0%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Penn Interactive

Penn Interactive

201 - 500 employees

🎲 Gambling

🎮 Gaming

🛍️ eCommerce

Gambling • Gaming • eCommerce

Penn Interactive is an interactive gaming company headquartered in Philadelphia, with additional offices in Greenfield, MA, and Cherry Hill, NJ. As the digital arm of PENN Entertainment, it manages various digital products, including the Barstool Sportsbook. The company is focused on fast-paced growth in the sports betting and online casino space, leveraging partnerships with Barstool Sports and theScore to offer a unique sports betting experience through mobile apps and retail locations.

📋 Description

• Build and operate infrastructure behind a large-scale sports betting and media platform • Own critical infrastructure across compute, networking, storage, and cloud services • Drive complex infrastructure migrations and projects across production environments and jurisdictions • Build and maintain platform tooling and automation using ArgoCD, Helm, GitHub Actions, release pipelines, and service onboarding workflows • Support development teams with infrastructure consulting, dependency resolution, architecture reviews, and platform-tool adoption • Design and improve Datadog observability, alerting, dashboards, and runbooks • Provide operational support and incident response through structured debugging and root cause analysis • Participate in on-call rotations • Mentor teammates and contribute to architecture decisions and continuous improvement

🎯 Requirements

• 5+ Years of Experience in a similar role (DevOps, Site Relatability Engineer) • Strong experience operating and troubleshooting Kubernetes in a production Linux environment (cluster lifecycle, networking, storage, scheduling) • Experience working with AWS, GCP, and/or on-premise environments • Proficiency in at least two of: Go, Python, Bash/Shell • Deep understanding of distributed systems, failure modes, networking fundamentals, capacity planning, and performance analysis • Experience with GitOps and CI/CD workflows (ArgoCD, Helm, GitHub Actions, or similar) • Experience with infrastructure-as-code (Terraform, Helm, or equivalent) • Track record of leading complex migrations or infrastructure projects with cross-team dependencies • Strong incident response and troubleshooting skills • Clear technical communication and documentation skills • Experience with service mesh technologies (Istio, Cilium) preferred • Familiarity with distributed storage systems (Ceph, or similar) preferred • Experience with bare-metal Kubernetes or Talos OS preferred • Exposure to regulated environments preferred • Experience with Datadog or comparable observability platforms at scale preferred • Familiarity with PostgreSQL, PgBouncer, or database migration tooling preferred

🏖️ Benefits

• Competitive compensation package • Comprehensive Benefits package • Fun, relaxed work environment • Education and conference reimbursements • Bonus eligibility for most non-sales positions • Best-in-class benefits with personalized physical, financial, and emotional support options

Apply Now

Similar Jobs

🕒 Yesterday

GE Vernova

10,000+ employees

💼 Consulting

📦 Logistics

🏭 Manufacturing

Senior Reliability Engineer improving embedded protection, control, and software products for utility grids. Leading reliability testing, failure analysis, and modernization initiatives for resilient energy systems.

🕒 Yesterday

GE Vernova

10,000+ employees

💼 Consulting

📦 Logistics

🏭 Manufacturing

Senior Reliability Engineer improving embedded grid automation reliability for utility-scale energy systems. Leading testing, failure analysis, KPIs, and modernization initiatives with utilities.

🕒 Yesterday

Valtech

5001 - 10000

💼 Consulting

📣 Marketing

☁️ SaaS

Site Reliability Expert leading observability and SRE for Valtech, an experience innovation company. Improving reliability across cloud-native, microservices-based environments.

🗣️🇫🇷 French Required

🕒 3 days ago

JFrog

1001 - 5000

💼 Consulting

🏥 Healthcare

📦 Logistics

Senior DevOps Engineer helping JFrog customers build CI/CD platforms using JFrog’s liquid software tools. Designing cloud-native pipelines and guiding customers, communities, and internal teams.

🕒 August 20

Mirantis

501 - 1000

💼 Consulting

🏥 Healthcare

📦 Logistics

Senior Site Reliability Engineer deploying Kubernetes-based AI infrastructure on NVIDIA-certified hardware for Mirantis. Ensuring reliable, secure, scalable cloud operations and customer delivery.