Site Reliability Engineer – Inference Infrastructure

Offre fantôme probable

🕒 il y a 7 mois

🇨🇦 Canada – Télétravail

⏰ Temps Plein

🟡 Intermédiaire

🟠 Senior

⛑ Ingénieur DevOps & SRE

👻 Score fantôme 63%

infoinfo

🗣️🇺🇸🇬🇧 Anglais requis

Postuler Maintenant
Trouver des Emplois à Distance Similaires

📊 Vérifiez votre score de CV pour ce poste

Améliorez vos chances d'obtenir un entretien en vérifiant votre score de CV avant de postuler.

Logo of Cohere

Cohere

11 - 50 employés

🤖 Intelligence artificielle

🏢 Entreprise

☁️ SaaS

Artificial Intelligence • Enterprise • SaaS

Cohere est une plateforme d’IA de premier plan qui fournit aux entreprises des modèles de langage avancés et un espace de travail (workspace) intégré conçu pour l’efficacité et la sécurité. Grâce à une famille de modèles génératifs et de retrieval hautes performances, Cohere permet aux organisations de rationaliser leurs workflows, de renforcer la sécurité des données et de révéler des insights dans de nombreux secteurs grâce à des capacités multilingues. Leur priorité accordée à des solutions d’IA sur mesure garantit la protection des données critiques tout en facilitant une intégration fluide aux processus opérationnels existants.

Description

• Build self-service systems that automate managing, deploying and operating services. • This includes our custom Kubernetes operators that support language model deployments. • Automate environment observability and resilience. Enable all developers to troubleshoot and resolve problems. • Take steps required to ensure we hit defined SLOs, including participation in an on-call rotation. • Build strong relationships with internal developers and influence the Infrastructure team’s roadmap based on their feedback. • Develop our team through knowledge sharing and an active review process.

🎯 Exigences

• 5+ years of engineering experience running production infrastructure at a large scale • Experience designing large, highly available distributed systems with Kubernetes, and GPU workloads on those clusters • Experience with Kubernetes dev and production coding and support • Experience with GCP, Azure, AWS, OCI, multi-cloud on-prem / hybrid serving • Experience in designing, deploying, supporting, and troubleshooting in complex Linux-based computing environments • Experience in compute/storage/network resource and cost management • Excellent collaboration and troubleshooting skills to build mission-critical systems, and ensure smooth operations and efficient teamwork • The grit and adaptability to solve complex technical challenges that evolve day to day • Familiarity with computational characteristics of accelerators (GPUs, TPUs, and/or custom accelerators), especially how they influence latency and throughput of inference. • Strong understanding or working experience with distributed systems. • Experience in Golang, C++ or other languages designed for high-performance scalable servers).

🏖️ Avantages

• An open and inclusive culture and work environment • Work closely with a team on the cutting edge of AI research • Weekly lunch stipend, in-office lunches & snacks • Full health and dental benefits, including a separate budget to take care of your mental health • 100% Parental Leave top-up for up to 6 months • Personal enrichment benefits towards arts and culture, fitness and well-being, quality time, and workspace improvement • Remote-flexible, offices in Toronto, New York, San Francisco, London and Paris, as well as a co-working stipend • 6 weeks of vacation (30 working days!)

Postuler Maintenant

Emplois Similaires

🕒 il y a 9 mois

Kong Inc.

201 - 500

💼 Conseil

📦 Logistique

🔌 API

Site Reliability Engineer responsible for operating and scaling Kong’s multi-region SaaS platform. Collaborating on infrastructure, automation, and ensuring service reliability across global regions.

🇨🇦 Canada – Télétravail

💰 €100 000 000 Series D en 2021-02

⏰ Temps Plein

🟠 Senior

⛑ Ingénieur DevOps & SRE

🗣️🇺🇸🇬🇧 Anglais requis

🕒 il y a 10 mois

Atolio

11 - 50

🤖 Intelligence artificielle

🏢 Entreprise

☁️ SaaS

Deployment Engineer working with engineering and client success teams at Atolio. Ensure efficient deployment of enterprise search platform in various environments.

🇨🇦 Canada – Télétravail

💵 CA$150 000 - CA$200 000 / an

⏰ Temps Plein

🟠 Senior

⛑ Ingénieur DevOps & SRE

🗣️🇺🇸🇬🇧 Anglais requis

🕒 il y a 11 mois

Veeva Systems

1001 - 5000

🏥 Santé

💼 Conseil

📦 Logistique

DevOps Engineer building scalable cloud and CI/CD infrastructure for Veeva Systems' life sciences SaaS. Focus on IaC, automation, Kubernetes, Terraform, and reliability.

🇨🇦 Canada – Télétravail

💵 CA$85 000 - CA$225 000 / an

⏰ Temps Plein

🟡 Intermédiaire

🟠 Senior

⛑ Ingénieur DevOps & SRE

🗣️🇺🇸🇬🇧 Anglais requis

🕒 il y a 11 mois

Veeva Systems

1001 - 5000

🏥 Santé

💼 Conseil

📦 Logistique

DevOps Engineer building scalable AWS infrastructure, CI/CD, and containerized deployments for Veeva's life sciences cloud; focuses on automation, reliability, and mentorship.

🇨🇦 Canada – Télétravail

💵 $85 000 - $225 000 / an

⏰ Temps Plein

🟡 Intermédiaire

🟠 Senior

⛑ Ingénieur DevOps & SRE

🗣️🇺🇸🇬🇧 Anglais requis

🕒 il y a 11 mois

Veeva Systems

1001 - 5000

🏥 Santé

💼 Conseil

📦 Logistique

DevOps Engineer building scalable cloud infrastructure at Veeva Systems. Ensuring reliable, automated delivery of SaaS products for life sciences customers.

🇨🇦 Canada – Télétravail

💵 $85 000 - $225 000 / an

⏰ Temps Plein

🟡 Intermédiaire

🟠 Senior

⛑ Ingénieur DevOps & SRE