
11 - 50 employés
🤖 Intelligence artificielle
🏢 Entreprise
☁️ SaaS
Artificial Intelligence • Enterprise • SaaS
Cohere est une plateforme d’IA de premier plan qui fournit aux entreprises des modèles de langage avancés et un espace de travail (workspace) intégré conçu pour l’efficacité et la sécurité. Grâce à une famille de modèles génératifs et de retrieval hautes performances, Cohere permet aux organisations de rationaliser leurs workflows, de renforcer la sécurité des données et de révéler des insights dans de nombreux secteurs grâce à des capacités multilingues. Leur priorité accordée à des solutions d’IA sur mesure garantit la protection des données critiques tout en facilitant une intégration fluide aux processus opérationnels existants.
🕒 il y a 5 mois
🗣️🇺🇸🇬🇧 Anglais requis
Améliorez vos chances d'obtenir un entretien en vérifiant votre score de CV avant de postuler.

11 - 50 employés
🤖 Intelligence artificielle
🏢 Entreprise
☁️ SaaS
Artificial Intelligence • Enterprise • SaaS
Cohere est une plateforme d’IA de premier plan qui fournit aux entreprises des modèles de langage avancés et un espace de travail (workspace) intégré conçu pour l’efficacité et la sécurité. Grâce à une famille de modèles génératifs et de retrieval hautes performances, Cohere permet aux organisations de rationaliser leurs workflows, de renforcer la sécurité des données et de révéler des insights dans de nombreux secteurs grâce à des capacités multilingues. Leur priorité accordée à des solutions d’IA sur mesure garantit la protection des données critiques tout en facilitant une intégration fluide aux processus opérationnels existants.
• Work across the inference stack to improve core performance metrics • Dive deep into model execution • Identify bottlenecks and develop innovative optimizations • Collaborate closely with modeling and systems teams • Experiment, measure, and ship improvements that accelerate inference • Build expertise in advanced performance techniques, including GPU/CUDA optimizations, kernel-level improvements, and model execution strategies for MoE and large-scale architectures
• 5+ years of experience writing high-performance, production-quality code • Strong programming skills in C++ or Python (Rust/Go also welcome) • Experience working with large language models and familiarity with the LLM inference ecosystem (e.g., vLLM, SGLang, etc.) • Ability to diagnose and resolve performance bottlenecks across the model execution stack • A strong bias for action — you ship fast, measure impact, and iterate • It’s a big plus if you have experience with GPU programming, CUDA, or low-level systems optimization • Language modeling with transformers (MoE, speculative decoding, KV-cache optimizations) • Scaling performance-critical distributed systems (e.g., computation, search, storage)
• An open and inclusive culture and work environment • Work closely with a team on the cutting edge of AI research • Weekly lunch stipend, in-office lunches & snacks • Full health and dental benefits, including a separate budget to take care of your mental health • 100% Parental Leave top-up for up to 6 months • Personal enrichment benefits towards arts and culture, fitness and well-being, quality time, and workspace improvement • Remote-flexible, offices in Toronto, New York, San Francisco, London and Paris, as well as a co-working stipend • 6 weeks of vacation (30 working days!)
Postuler Maintenant🕒 il y a 5 mois
Chief Development Officer leading revenue growth strategies for colorectal cancer nonprofit. Engaging in high-level fundraising and team leadership to significantly increase funding.
🗣️🇺🇸🇬🇧 Anglais requis
🕒 il y a 5 mois
Chief Business Officer leading global commercial function of an expanding neuroscience startup. Driving strategy and partnerships to position the organization as a leading CRO partner in biotech and pharma.
🗣️🇺🇸🇬🇧 Anglais requis