Senior Site Reliability Engineer – Kubernetes

Emploi pas sur LinkedIn

🕒 il y a 14 jours

🇨🇦 Canada – Télétravail

⏰ Temps Plein

🟠 Senior

⛑ Ingénieur DevOps & SRE

👻 Score fantôme 15%

infoinfo

🗣️🇺🇸🇬🇧 Anglais requis

Postuler Maintenant
Trouver des Emplois à Distance Similaires

📊 Vérifiez votre score de CV pour ce poste

Améliorez vos chances d'obtenir un entretien en vérifiant votre score de CV avant de postuler.

Logo of Software Mind

Software Mind

1001 - 5000 employés

Fondée en 1999

🤖 Intelligence artificielle

☁️ SaaS

📡 Télécommunications

💰 Private Equity Round en 2020-12

Artificial Intelligence • SaaS • Telecommunications

Software Mind est une entreprise technologique spécialisée dans le développement logiciel et les services de transformation numérique. En mettant l'accent sur les solutions IA et cloud, l'entreprise propose une large gamme de services, incluant le développement de logiciels sur mesure, le développement d'applications mobiles et le conseil en cloud. Software Mind dessert divers secteurs tels que les services financiers, les télécommunications, la biotechnologie et les médias, fournissant des solutions personnalisées pour accélérer les transformations numériques et la croissance des entreprises à l'échelle mondiale.

Description

• Support the deployment, operation, and reliability of production services running on Kubernetes • Monitor service health and investigate production incidents across distributed applications • Participate in on-call support, incident response, root cause analysis, postmortems, and reliability improvements • Troubleshoot application runtime, networking, and service-to-service issues in collaboration with engineering teams • Support CI/CD, GitOps-based deployments, observability, and production monitoring • Work within a client-directed backlog and established priorities

🎯 Exigences

• 5+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, Production Engineering, or a closely related role • Strong recent hands-on experience supporting Kubernetes-based production services • 3+ years of hands-on production Kubernetes experience strongly preferred • Kubernetes production operations, including deployment, scaling, rollout/rollback, resource tuning, and service-to-service troubleshooting • Strong production incident response experience, including on-call, runbooks, postmortems, and paging hygiene • Splunk experience for log aggregation, search, and production troubleshooting • Prometheus and Grafana experience, specifically building alert rules and dashboards • CI/CD and infrastructure-as-code for containerized deployments, including Helm and GitOps tools such as ArgoCD or Flux • Strong Linux and networking fundamentals, including DNS, load balancing, TCP/HTTP, HTTP/2, and Kubernetes networking • Node.js production troubleshooting, including heap snapshots, CPU profiles, event-loop blocking, memory growth, worker/process isolation, and V8 isolates or similar runtime models • JVM/Java production troubleshooting, including GC log analysis, thread dump analysis, JVM tuning, and Java service latency investigation • In-memory cache experience with Redis/Valkey, including key design, TTL/eviction tuning, and cache invalidation • Service-to-service authentication experience, including mTLS, certificate rotation, certificate format conversion, and JWT-based service authentication • Web Components/Lit experience is nice to have • Server-side rendering or isomorphic runtime experience is nice to have • Canary rollout/multi-version production operations is nice to have • Distributed tracing and request-context correlation is nice to have • KEDA or event-driven autoscaling experience is nice to have • Experience with enterprise platform integration layers is nice to have

🏖️ Avantages

• Competitive salary and laptop • Professional development and training opportunities • Work with cutting-edge cloud and container technologies • Flexible work arrangements and collaborative team environment • Impact on organization-wide digital transformation initiatives

Postuler Maintenant

Emplois Similaires

🕒 il y a 16 jours

Autodesk

10 000+ employés

🏗️ Construction

🏭 Fabrication

💼 Conseil

Senior DevOps Developer building reliable AWS, Kubernetes, and MongoDB services for Autodesk Construction Solutions. Improving automation, observability, security, disaster recovery, and production reliability for construction software customers.

🇨🇦 Canada – Télétravail

💵 $107 000 - $157 300 / an

⏰ Temps Plein

🟠 Senior

⛑ Ingénieur DevOps & SRE

🗣️🇺🇸🇬🇧 Anglais requis

🕒 il y a 16 jours

High Tech Genesis

51 - 200

📦 Logistique

📣 Marketing

🏭 Fabrication

Cloud DevOps Engineer building AWS infrastructure, CI/CD pipelines, and automation for High Tech Genesis. Supporting containers, event-driven systems, service mesh, and observability.

🇨🇦 Canada – Télétravail

💵 CA$65 - CA$70 / heure

⏰ Temps Plein

🟡 Intermédiaire

🟠 Senior

⛑ Ingénieur DevOps & SRE

🗣️🇺🇸🇬🇧 Anglais requis

🕒 il y a 20 jours

Yelp

1001 - 5000

🍽️ Alimentation et boissons

🏨 Hôtellerie

📣 Marketing

Site Reliability Engineer operating Yelp’s Kafka and Flink streaming platform across Canada. Automating cluster management, scaling, upgrades, migrations, and incident recovery for real-time data systems.

🇨🇦 Canada – Télétravail

💵 $135 000 - $185 000 / an

⏰ Temps Plein

🟡 Intermédiaire

🟠 Senior

⛑ Ingénieur DevOps & SRE

🗣️🇺🇸🇬🇧 Anglais requis

🕒 il y a 21 jours

Yelp

1001 - 5000

🍽️ Alimentation et boissons

🏨 Hôtellerie

📣 Marketing

Site Reliability Engineer operating Yelp’s Kafka and Flink streaming infrastructure across Canada. Automating cluster operations, upgrades, scaling, and incident recovery for real-time data systems.

🇨🇦 Canada – Télétravail

💵 $135 000 - $185 000 / an

⏰ Temps Plein

🟡 Intermédiaire

🟠 Senior

⛑ Ingénieur DevOps & SRE

🗣️🇺🇸🇬🇧 Anglais requis

🕒 il y a 27 jours

Mirantis

501 - 1000

💼 Conseil

🏥 Santé

📦 Logistique

Senior Site Reliability Engineer at Mirantis, contributing to cloud-based AI solutions using Kubernetes. Focused on deploying AI infrastructure and ensuring system reliability and performance.

🇨🇦 Canada – Télétravail

⏰ Temps Plein

🟠 Senior

⛑ Ingénieur DevOps & SRE

🗣️🇺🇸🇬🇧 Anglais requis