Site Reliability Engineer – AI

Job not on LinkedIn

🕒 April 15

🇵🇱 Poland – Remote

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

👻 Ghost score 36%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Madiff

Madiff

51 - 200 employees

Founded 2015

💼 Consulting

📦 Logistics

🏥 Healthcare

Consulting • Logistics • Healthcare

Madiff is a technology and consulting company specializing in IT services and digital innovation. They provide a range of solutions, including software development, agile team services, and IT consulting across various industries. With a focus on delivering high-quality, multi-skilled remote teams, Madiff aims to enhance business performance and support clients in navigating the complexities of their respective markets.

📋 Description

• Build and maintain central monitoring and alerting layer for AI applications and pipelines • Define and implement SLIs, alerts, and operational dashboards • Manage incidents including triage, coordination, root cause analysis, and prevention • Standardise telemetry across systems including latency, throughput, and failures • Optimise CI CD pipelines and introduce quality gates for reliability • Work closely with engineering teams to reduce recurring issues and improve stability

🎯 Requirements

• Minimum 5+ years of experience in SRE, Platform, or Production Engineering • Strong hands on experience with Kubernetes and production environments • Experience with Azure and Azure DevOps • Experience with monitoring tools such as Datadog • Strong understanding of incident management and root cause analysis • Ability to build practical monitoring and alerting systems • Nice to have: Experience with AI or LLM pipelines • Nice to have: Experience building monitoring platforms across multiple systems • Nice to have: Experience with Grafana • Nice to have: Experience working in large scale or distributed environments

🏖️ Benefits

• Solid, competitive salary • Work in a multinational environment on international projects • Comprehensive healthcare • Long-term B2B contract with a stable project pipeline • Work model: fully remote

Apply Now

Similar Jobs

🕒 April 2

Akamai Technologies

5001 - 10000

🔒 Cybersecurity

Senior SRE responsible for automation, architecture decisions, and managing AI workloads for Akamai. Collaborating with product teams on reliability and operational readiness.

Distributed Systems

Grafana

Kubernetes

Prometheus

Python

Terraform

Go

🕒 April 2

Akamai Technologies

5001 - 10000

🔒 Cybersecurity

🏢 Enterprise

📱 Media

Site Reliability Engineer II at Akamai focusing on AI platforms, automation, and system reliability. Collaborating with teams for incident response and workload monitoring.

Distributed Systems

Grafana

Kubernetes

Linux

Prometheus

Python

SaltStack

Terraform

Go

🕒 April 2

Akamai Technologies

5001 - 10000

🔒 Cybersecurity

🏢 Enterprise

📱 Media

Senior SRE responsible for reliability workstreams and automation for Akamai's serverless inference platform. Engage in architecture decisions and expertise in GPU infrastructure and AI inference workloads.

Distributed Systems

Grafana

Kubernetes

Prometheus

Python

Terraform

Go

🕒 April 2

Akamai Technologies

5001 - 10000

🔒 Cybersecurity

🏢 Enterprise

📱 Media

Lead reliability workstreams for Akamai's serverless inference platform as Senior II SRE. Responsible for observability, incident management integration, and mentoring engineering teams.

Distributed Systems

Kubernetes

Python

Go

🕒 April 1

Base.com

51 - 200

📦 Logistics

📣 Marketing

💼 Consulting

Senior AWS DevOps Engineer managing cloud infrastructure on AWS. Ensuring stable service operation while handling Unix/Linux servers and incident response.

🗣️🇵🇱 Polish Required

AWS

Linux

Unix