Lead Software Platform Engineer, MLOps

🕒 il y a 2 jours

🏄 California, Massachusetts – Distant

info

💵 $200 000 - $270 000 / an

⏰ Temps Plein

🟠 Senior

🏗️ Ingénieur Plateforme

🦅 Parrain de Visa H1B

info

🗣️🇺🇸🇬🇧 Anglais requis

Postuler Maintenant
Trouver des Emplois à Distance Similaires

📊 Vérifiez votre score de CV pour ce poste

Améliorez vos chances d'obtenir un entretien en vérifiant votre score de CV avant de postuler.

Logo of TetraScience

TetraScience

51 - 200 employés

Fondée en 2015

🏥 Santé

💼 Conseil

📦 Logistique

Healthcare • Consulting • Logistics

TetraScience est une entreprise dédiée à la transformation des données scientifiques brutes en ensembles de données natifs pour l'IA, destinés à des applications scientifiques avancées. En collaborant étroitement avec les principales entreprises biopharmaceutiques, TetraScience améliore la productivité, accélère les insights et assure l'intégrité des données tout au long de la chaîne de valeur scientifique. Leur plateforme propose des solutions pour la gestion des données de laboratoire de nouvelle génération, des résultats scientifiques pilotés par l'IA et la conformité avec les normes industrielles. En tant que première entreprise à fournir un cloud de données et d'IA conçu spécifiquement pour la science, TetraScience permet à ses clients de libérer, unifier et transformer leurs données, surmontant les silos de données traditionnels et boostant la productivité scientifique grâce à une infrastructure flexible, ouverte et collaborative.

Description

• Own the technical architecture of the AI/ML platform and its customer-facing service and API surface. • Own the end-to-end model and prompt lifecycle across Databricks MLflow and AWS Bedrock, including registration, versioning, staged promotion, rollback, and multi-model serving. • Design inference infrastructure for real-time and batch AI workloads, including routing, batching, caching, concurrency control, GPU capacity planning, large binary inputs, and graceful degradation. • Integrate AI models and LLMs into production systems using RAG, tool and function calling, MCP-based tooling, and agent runtimes. • Design platform security, including guardrails, prompt-injection and tool-abuse defenses, PII and PHI handling, and tenant data boundaries. • Build AI evaluation and quality infrastructure, including offline and online evaluation harnesses, golden datasets, CI regression gates, A/B and shadow deployment, and drift and hallucination detection. • Establish monitoring, alerting, logging, distributed tracing, and SLI/SLO/SLA practices for the AI platform. • Design reproducibility and lineage for validated pharma environments, including versioned data, code, prompts, and model artifacts and audit trails. • Contribute to infrastructure-as-code and deployment automation using CloudFormation and AWS CDK. • Own production readiness with Applied AI, data engineering, and platform teams, including performance, reliability, cost efficiency, incident response, and runbooks. • Act as SME and design authority, lead design reviews, write reference architectures and technical documentation, and mentor engineers. • Evaluate emerging AI infrastructure frameworks, serving runtimes, model providers, and data types, making build-versus-buy decisions.

🎯 Exigences

• 10+ years of professional experience in software engineering and infrastructure engineering designing, building, and scaling distributed cloud-native production systems. • Proven technical leadership or architecture experience with accountability for system design, scalability, performance, and cost optimization. • Experience designing security into multi-tenant platforms, including tenant authorization boundaries and PII/PHI handling. • Awareness of LLM-specific risks such as prompt injection and tool abuse. • Extensive production experience building AI/ML infrastructure as a multi-tenant product for external users. • Deep hands-on experience taking LLM-based systems to production, including RAG, retrieval and embedding design, prompt and model versioning, and tool or function calling. • Expert-level TypeScript and Python coding skills for robust APIs and backend services. • Production experience with model registries and serving stacks, ideally Databricks MLflow. • Experience with AI evaluation release gates, regression gates, and drift or quality monitoring. • Proficiency in API-first design, REST, and OpenAPI. • Solid AWS and Docker knowledge. • Familiarity with CloudFormation or AWS CDK, CI/CD pipelines, and deployment automation. • Experience defining observability and SLI/SLO/SLA practices, including monitoring, alerting, and distributed tracing. • Ability to communicate with customers and cross-functional teams, influence technical direction, and mentor engineers. • Nice to have: emerging LLM frameworks, production agentic orchestration, MCP, LLM cost monitoring, multimodal inputs, fine-tuning or model optimization, regulated or validated environments, and scientific or life sciences domains. • Visa sponsorship is not currently provided for this position.

🏖️ Avantages

• 100% employer-paid benefits for all eligible employees and immediate family members • Unlimited paid time off (PTO) • 401K • Flexible working arrangements - Remote work • Company paid Life Insurance, LTD/STD • A culture of continuous improvement where you can grow your career and get coaching

Postuler Maintenant

Emplois Similaires

🕒 il y a 2 jours

CrowdStrike

5001 - 10000

🔒 Cybersecurity

☁️ SaaS

🤖 Intelligence artificielle

AI Platform Engineer scaling AWS/GCP infrastructure for CrowdStrike’s internal cybersecurity agents. Building secure, observable, production-ready platforms from rapid AI prototypes.

🇺🇸 États-Unis – Télétravail

💵 $140 000 - $215 000 / an

⏰ Temps Plein

🟡 Intermédiaire

🟠 Senior

🏗️ Ingénieur Plateforme

🦅 Parrain de Visa H1B

info

🗣️🇺🇸🇬🇧 Anglais requis

🕒 il y a 2 jours

Conduent

10 000+ employés

🏥 Santé

📦 Logistique

💼 Conseil

Azure Linux & PaaS Platform Engineer supporting Conduent’s mission-critical services and solutions. Managing Azure workloads, observability, networking, security, troubleshooting, and modernization.

🇺🇸 États-Unis – Télétravail

💵 $85 470 - $111 000 / an

💰 Venture Round en 2009-01

⏰ Temps Plein

🟡 Intermédiaire

🟠 Senior

🏗️ Ingénieur Plateforme

🦅 Parrain de Visa H1B

info

🗣️🇺🇸🇬🇧 Anglais requis

🕒 il y a 2 jours

ZenBusiness

201 - 500

⚖️ Juridique

💼 Conseil

📣 Marketing

Senior Software Engineer building AI-powered developer tooling, CI/CD, and Kubernetes platforms. Improving delivery speed, reliability, security, and developer adoption for business-launch services.

🗣️🇺🇸🇬🇧 Anglais requis

🕒 il y a 5 jours

Cornelis Networks

51 - 200

🤖 Intelligence artificielle

🔧 Matériel

🏢 Entreprise

AI Platform Engineer for Cornelis Networks, designing AI systems for advanced networking solutions. Focus on building private, secure AI platforms to optimize engineering work.

🇺🇸 États-Unis – Télétravail

💰 €29 000 000 Series B en 2022-11

⏰ Temps Plein

🟡 Intermédiaire

🟠 Senior

🏗️ Ingénieur Plateforme

🦅 Parrain de Visa H1B

info

🗣️🇺🇸🇬🇧 Anglais requis

🕒 il y a 5 jours

Aderant

501 - 1000

⚖️ Juridique

💼 Conseil

☁️ SaaS

Solutions Platform Engineer for Aderant's demonstration platform, troubleshooting technical issues and supporting deployments. Collaborating with multiple teams to maintain system functionality and user access.

🗣️🇺🇸🇬🇧 Anglais requis

Azure

Cloud

MS SQL Server

SQL