Engineering Manager – Site Reliability

Job not on LinkedIn

🔥 1 minute ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Pliant

Pliant

201 - 500 employees

Founded 2020

💳 Fintech

☁️ SaaS

🤝 B2B

💰 $40M Series B - Pliant on 2025-04

Fintech • SaaS • B2B

Pliant is a credit card platform that enables businesses, banks, and fintechs to issue and manage credit cards, automate payments, and optimize spend and accounting workflows. Pliant offers Payment Apps, a Pro API for card issuance and automation, Cards-as-a-Service (CaaS) and Banking-as-a-Service (BaaS) to embed or white-label card programs, plus global accounts, FX and transfers, credit & financing, and compliance and enablement services. It serves corporations, e-commerce, resellers, SaaS companies, travel and marketing agencies, and banks with features like real-time monitoring, spend controls, receipt management, integrations with accounting and expense systems, single-use virtual cards, and API-driven automation. Pliant holds e-money licensing in the EU, issues cards in the UK, is PCI DSS and ISO/IEC 27001 certified, and focuses on B2B payment optimization and embedded finance solutions.

📋 Description

• Define the framework for teams to set SLOs and error budgets, educating and supporting product teams • Own blameless post-mortems and root-cause fixes • Implement production readiness reviews for new releases • Improve Datadog observability coverage, including alerts, dashboards, and on-call pages • Hire and build the Site Reliability team from the ground up • Build the on-call rotation and incident management process • Establish SLOs and an incident review process • Integrate reliability into the software development lifecycle and reduce repeat incidents

🎯 Requirements

• 7–10 years of engineering experience, including at least 3 years directly managing engineers • Track record of hiring and developing engineers, including levelling or promoting team members • Hands-on production or reliability engineering background • Prior experience carrying a pager; this is not a first management role • Strong AWS and Terraform experience • Comfortable working inside a managed infrastructure-as-code pipeline • Experience building or running an on-call rotation and incident management process • Strong platform observability experience • Clear communication for a technical, cross-team audience • Track record of introducing reliability practices into product engineering teams • Proficiency with AI-assisted development tools such as Claude Code and Cursor • Ability to rigorously review AI-written pull requests • Familiarity with PCI DSS, SOC 2, and ISO 27001 environments is relevant to the stack/context

🏖️ Benefits

• Attractive remuneration • Choice of preferred OS: Windows or Mac • Flat hierarchy and transparent communication in a relaxed, professional atmosphere • Opportunity to develop your talent in a dynamic team with ambitious goals • Flexibility and possibility to work remotely • Pliant Card with monthly credit to explore the product and enjoy food with colleagues

Apply Now

Similar Jobs

🕒 5 days ago

UNITYTECH CONSULTING

11 - 50

💼 Consulting

Developer Support Engineer leading technical onboarding and solution design for key enterprise customers. Optimizing VPC deployments while collaborating with internal teams for enhancements and performance.

🕒 6 days ago

PatSnap

501 - 1000

💼 Consulting

🏥 Healthcare

📦 Logistics

Site Reliability Engineering Leader at PatSnap, leading the SRE team ensuring reliability for a global SaaS platform. Overseeing strategy, automation, and team development in cloud technologies.

🕒 July 28

TwinStream

51 - 200

🎖️ Defense

💼 Consulting

📦 Logistics

DevOps Engineer maintaining and deploying cross-domain systems using Docker and AMQP architecture for TwinStream clients. Collaborating with teams and ensuring system performance and availability.

🕒 July 28

RTX

10,000+ employees

🚀 Aerospace

🎖️ Defense

🏭 Manufacturing

Principal Site Reliability Engineer managing AWS infrastructures for Collins Aerospace. Delivering B2B products and ensuring service availability with scalable solutions in aviation technology.

🕒 July 27

Cognativ

11 - 50

💼 Consulting

🥽 AR/VR

🤖 Artificial Intelligence

Senior Site Reliability Engineer managing reliability for a distributed, camera-based video monitoring and AI alerting platform. Focusing on operational health, service objectives, and incident response.