Lead ML Ops/DevOps Engineer – AI Engineering

🕒 vor 1 Monat

🇺🇸 Vereinigte Staaten – Remote

💵 $140.000 - $220.000 / Jahr

⏰ Vollzeit

🟠 Senior

⛑ DevOps- und Site Reliability Engineer (SRE)

👻 Geisterscore 2%

infoinfo

🗣️🇺🇸🇬🇧 Englisch erforderlich

Jetzt Bewerben
Ähnliche Remote-Jobs finden

📊 Überprüfen Sie Ihre Lebenslauf-Bewertung für diese Stelle

Verbessern Sie Ihre Chancen auf ein Vorstellungsgespräch, indem Sie Ihre Lebenslauf-Bewertung vor der Bewerbung überprüfen.

Logo of FICO

FICO

1001 - 5000 Mitarbeiter

Gegründet 1956

💼 Beratung

🛡️ Versicherung

🏥 Gesundheitswesen

Consulting • Insurance • Healthcare

FICO ist ein führendes Analytics- und Softwareunternehmen, bekannt für den FICO® Score – ein von Kreditgebern weit verbreitetes Instrument zur Beurteilung des Kreditrisikos. Das Unternehmen bietet eine umfassende Plattform, die Daten, KI und Machine Learning nutzt, um intelligente Entscheidungsfindung und Customer Engagement in verschiedensten Branchen zu ermöglichen. Die Lösungen von FICO umfassen Fraud Detection, Credit Scoring und Customer Lifecycle Management und sind damit für Sektoren wie Finanzdienstleistungen und Telekommunikation essenziell. Die innovativen Produkte unterstützen Unternehmen dabei, Ergebnisse durch Echtzeit-Analytics, Business Composability und Scenario Management zu optimieren.

Beschreibung

• Design, build, and maintain scalable, resilient data and ML pipelines, infrastructure, and workflows using tools such as Terraform, GitHub Actions, ArgoCD, Helm, and others. • Automate infrastructure provisioning and configuration management using cloud-native services (preferably AWS) with tools like Terraform, CloudFormation. • Design, containerize, and manage Kubernetes (EKS) clusters and/or ECS environments in AWS. • Collaborate with development teams to optimize performance, deployment, and cost. • Partner with DevOps and SRE teams to ensure high availability, observability, scalability, and security of the data and ML infrastructure. • Work closely with Data Scientists and ML Engineers to operationalize machine learning models, including building CI/CD pipelines for model training, validation, and deployment. • Implement observability for data pipelines and ML services using tools like Prometheus, Grafana, Datadog, or similar. • Develop and maintain automated pipelines for model retraining, monitoring drift, and versioning in production. • Support experimentation and prototyping in areas such as Machine Learning and Generative AI, transitioning successful prototypes into production systems. • Ensure cloud infrastructure is secure, compliant, and cost-efficient, following best practices in governance, identity, and access management.

🎯 Anforderungen

• 8+ years of experience in DataOps, MLOps, or related fields, with 3+ years focused on ML model operationalization and workflow automation. • Proficient in AWS services including EC2, S3, IAM, ACM, Route 53, CloudWatch, EKS, and ECS. • Experience with infrastructure as code (IaC) tools such as Terraform, CloudFormation, and Helm. • Familiarity with CI/CD for ML pipelines, GitOps practices, and tools like GitHub Actions, Jenkins, or Argo Workflows. • Strong scripting and automation skills using Python, or GitHub workflows. • Solid understanding of observability and monitoring tools (e.g., Prometheus, Grafana, Datadog, or OpenTelemetry). • Solid understanding of security best practices for cloud and Kubernetes environments, including secrets management, identity & access control, and policy enforcement. • Strong understanding with data governance, lineage, and metadata management is a plus. • Excellent collaboration and communication skills, with a proven ability to work effectively in cross-functional, globally distributed teams. • A bachelor’s degree in computer sciences, or a related discipline, or equivalent hands-on industry experience.

🏖️ Vorteile

• An inclusive culture strongly reflecting our core values: Act Like an Owner, Delight Our Customers and Earn the Respect of Others. • The opportunity to make an impact and develop professionally by leveraging your unique strengths and participating in valuable learning experiences. • Highly competitive compensation, benefits and rewards programs that encourage you to bring your best every day and be recognized for doing so. • An engaging, people-first work environment offering work/life balance, employee resource groups, and social events to promote interaction and camaraderie.

Jetzt Bewerben

Ähnliche Jobs

🕒 vor 1 Monat

ICF

5001 - 10000

🏥 Gesundheitswesen

📦 Logistik

📣 Marketing

DevOps Engineer building best in class health care reporting service and automating deployment processes using AWS. Collaborating with cross-functional teams to enhance infrastructure and CI/CD efficiencies.

🇺🇸 Vereinigte Staaten – Remote

💵 $108.476 - $184.409 / Jahr

💰 €30.000.000 Grant im 2021-03

⏰ Vollzeit

🟠 Senior

⛑ DevOps- und Site Reliability Engineer (SRE)

🦅 H1B-Visum-Sponsor

infoinfo

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 1 Monat

Akamai Technologies

5001 - 10000

🔒 Cybersecurity

Senior II Site Reliability Engineer ensuring performance and reliability of Akamai's digital platform. Leading technical teams to address complex content delivery challenges.

🇺🇸 Vereinigte Staaten – Remote

💵 $146.400 - $263.600 / Jahr

💰 Post-IPO Equity im 2001-07

⏰ Vollzeit

🟠 Senior

⛑ DevOps- und Site Reliability Engineer (SRE)

🦅 H1B-Visum-Sponsor

infoinfo

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 1 Monat

Horizon Industries, Limited

201 - 500

💼 Beratung

☁️ SaaS

🔒 Cybersecurity

DevSecOps Engineer facilitating software development solutions for Horizon Industries. Transitioning existing solutions to micro-services and managing CI/CD pipelines.

🇺🇸 Vereinigte Staaten – Remote

⏰ Vollzeit

🟡 Mittelstufe

🟠 Senior

⛑ DevOps- und Site Reliability Engineer (SRE)

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 1 Monat

Filevine

201 - 500

☁️ SaaS

⚖️ Rechtswesen

🤖 Künstliche Intelligenz

Senior Site Reliability Engineer at Filevine providing observability excellence in software development. Collaborating on legal AI systems to enhance operational reliability and performance.

🇺🇸 Vereinigte Staaten – Remote

💵 $175.000 - $195.000 / Jahr

💰 €108.000.000 Series D im 2022-04

⏰ Vollzeit

🟠 Senior

⛑ DevOps- und Site Reliability Engineer (SRE)

🦅 H1B-Visum-Sponsor

infoinfo

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 1 Monat

Coterie

11 - 50

👥 B2C

🛍️ eCommerce

🛒 Einzelhandel

Senior Site Reliability Engineer joining Coterie to maintain reliable, scalable infrastructure to support high-quality software. Collaborating with teams and enhancing observability and incident response capabilities.

🗣️🇺🇸🇬🇧 Englisch erforderlich