Senior SRE

🕒 il y a 10 jours

🇨🇦 Canada – Télétravail

⏰ Temps Plein

🟠 Senior

⛑ Ingénieur DevOps & SRE

👻 Score fantôme 10%

infoinfo

🗣️🇺🇸🇬🇧 Anglais requis

Postuler Maintenant
Trouver des Emplois à Distance Similaires

📊 Vérifiez votre score de CV pour ce poste

Améliorez vos chances d'obtenir un entretien en vérifiant votre score de CV avant de postuler.

Logo of CloudFactory

CloudFactory

1001 - 5000 employés

Fondée en 2010

💼 Conseil

📦 Logistique

📣 Marketing

Consulting • Logistics • Marketing

CloudFactory est une entreprise qui propose une plateforme de données d'IA permettant d'accélérer la transition des concepts d'IA vers des solutions entièrement opérationnelles. Leur plateforme permet la création et la préparation de jeux de données structurés et de haute qualité visant à améliorer la précision des modèles et à accélérer le déploiement. Les services de CloudFactory incluent GenAI pour améliorer les performances des modèles LLM de base, la supervision des modèles pour auditer et gérer les modèles d'IA, et des services professionnels guidant les projets des étapes de preuve de valeur et de MVP jusqu'à la production complète. L'entreprise met fortement l'accent sur l'optimisation de l'inférence et sur la synergie entre l'intelligence humaine et machine pour garantir un déploiement fiable et digne de confiance des modèles d'IA, en se concentrant sur la fourniture de résultats concrets et réels pour les entreprises. Ils privilégient également la sécurité des données et le respect des normes de l'industrie telles que ISO 9001:2015, ISO 27001, SOC 2, HIPAA et GDPR.

Description

• Design and implement new core infrastructure components with a high degree of autonomy • Optimize and improve deployment pipelines, environment provisioning, and high-throughput batch jobs • Use Infrastructure as Code tools such as Terraform to manage and scale complex infrastructure • Develop CI/CD pipelines for automated build, test, deployment, and monitoring processes • Create and manage multi-step CI/CD pipelines, including environment setup and artifact handling • Support the reliability, availability, and performance of production systems • Apply site-reliability practices across infrastructure • Set up monitoring, alerting, and observability tooling • Collaborate with software engineers, product, and business stakeholders on infrastructure and deployment systems • Communicate complex technical issues clearly to technical and non-technical stakeholders

🎯 Exigences

• 5+ years of experience building and operating infrastructure in production environments • Fluent in Python with strong experience writing production-ready code • Experience with Docker and Kubernetes • Knowledge of cloud platforms such as GCP or AWS • Experience using Infrastructure as Code tools such as Terraform • Experience using CI/CD platforms to automate build, test, and deployment pipelines • Comfortable applying site-reliability principles including availability, observability, and automation • Degree in Computer Science, Engineering, or another quantitative or computational field, or equivalent practical experience • Familiarity with Prometheus or Grafana preferred • Experience with configuration management tools such as Ansible, Chef, or Puppet preferred • Exposure to multi-cloud or hybrid-cloud environments preferred

🏖️ Avantages

• Great Mission and Culture • Meaningful Work • Market competitive salary • Quarterly variable compensation • Hybrid Working Model • Comprehensive medical cover • Group life insurance • Personal development and growth opportunities

Postuler Maintenant

Emplois Similaires

🕒 il y a 10 jours

InnoData

2 - 10

🤝 B2B

💼 Conseil

🌍 Impact social

Application Reliability Engineer supporting Innodata’s Google Cloud enterprise applications. Restoring production services, managing deployments, and enhancing microservices for a global AI data engineering company.

🇨🇦 Canada – Télétravail

💵 $80 000 - $150 000 / an

⏰ Temps Plein

🟡 Intermédiaire

🟠 Senior

⛑ Ingénieur DevOps & SRE

🗣️🇺🇸🇬🇧 Anglais requis

🕒 il y a 13 jours

WinAir

51 - 200

📦 Logistique

💼 Conseil

🏭 Fabrication

DevOps Specialist automating CI/CD and infrastructure for WinAir’s aviation maintenance software. Improving Jenkins, Ansible, Linux environments, deployments, and monitoring across development and production systems.

🇨🇦 Canada – Télétravail

💵 $54 000 - $76 000 / an

⏰ Temps Plein

🟡 Intermédiaire

🟠 Senior

⛑ Ingénieur DevOps & SRE

🗣️🇺🇸🇬🇧 Anglais requis

🕒 il y a 14 jours

Software Mind

1001 - 5000

🤖 Intelligence artificielle

☁️ SaaS

📡 Télécommunications

Senior SRE maintaining Kubernetes-based UI and AI service reliability for an enterprise cloud software company. Managing incidents, observability, deployments, and runtime troubleshooting in production.

🇨🇦 Canada – Télétravail

💰 Private Equity Round en 2020-12

⏰ Temps Plein

🟠 Senior

⛑ Ingénieur DevOps & SRE

🗣️🇺🇸🇬🇧 Anglais requis

🕒 il y a 14 jours

Software Mind

1001 - 5000

🤖 Intelligence artificielle

☁️ SaaS

📡 Télécommunications

Senior SRE supporting Kubernetes production reliability for an enterprise cloud software company. Troubleshooting distributed services, observability, incidents, CI/CD, Node.js, and JVM/Java runtimes.

🇨🇦 Canada – Télétravail

💰 Private Equity Round en 2020-12

⏰ Temps Plein

🟠 Senior

⛑ Ingénieur DevOps & SRE

🗣️🇺🇸🇬🇧 Anglais requis

🕒 il y a 15 jours

Software Mind

1001 - 5000

🤖 Intelligence artificielle

☁️ SaaS

📡 Télécommunications

Site Reliability Engineer operating Kubernetes UI services for a leading enterprise cloud software company. Monitoring reliability, troubleshooting incidents, and supporting production operations.

🇨🇦 Canada – Télétravail

💰 Private Equity Round en 2020-12

⏰ Temps Plein

🟡 Intermédiaire

🟠 Senior

⛑ Ingénieur DevOps & SRE

🗣️🇺🇸🇬🇧 Anglais requis