Search Remote Jobs

Head of SRE

đŸ”„ 17 minutes ago

đŸ‡ȘđŸ‡ș Europe – Remote

⏰ Full Time

🔮 Lead

⛑ DevOps & Site Reliability Engineer (SRE)

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Wand AI

Wand AI

51 - 200 employees

Founded 2022

đŸ€– Artificial Intelligence

🏱 Enterprise

☁ SaaS

Artificial Intelligence ‱ Enterprise ‱ SaaS

Wand AI is an enterprise software company (Wand Synthesis AI Inc. ) that builds the “Agentic Labor Infrastructure” to enable governments and large enterprises to create, manage, and scale hybrid workforces composed of humans and autonomous AI agents. Their platform (branded Wand OS / Agentic Workforce Technology) provides management, oversight, interoperability across systems, security options (SOC2-ready, on‑premise/private cloud/hosted), dashboards, decision tracking, and tools for deploying and governing agentic workflows at scale. Wand positions itself as a B2B/enterprise provider that turns AI into operational labor for regulated and large-scale organizations.

📋 Description

‱ Own and lead all SRE-related strategy, standards, and execution. ‱ Embed SRE culture and operational excellence across engineering teams. ‱ Review the current infrastructure and operational model; redesign and rebuild where needed. ‱ Architect, deploy, and maintain scalable, secure production environments. ‱ Define and implement SLIs, SLOs, and uptime targets. ‱ Establish robust monitoring, alerting, and observability practices. ‱ Design and implement incident management, RCA and postmortem processes. ‱ Build and manage sustainable on-call frameworks and escalation models. ‱ Automate the software delivery lifecycle to improve release predictability and safety. ‱ Create reproducible environments and IaaC provisioning templates. ‱ Improve system performance, availability, and reliability. ‱ Support and productionise data platforms and ML workloads. ‱ Partner closely with QA and Engineering leadership to improve release quality and stability. ‱ Ensure infrastructure meets enterprise-grade security and regulatory requirements. ‱ Hire, manage, and mentor a team of SRE engineers.

🎯 Requirements

‱ Proven hands-on experience in Site Reliability Engineering, Production Engineering, or a similar role. ‱ Strong hands-on expertise in cloud infrastructure (AWS or Azure preferred), IaaC (Terraform) and Kubernetes. ‱ Experience building or maturing SRE practices within an organisation. ‱ Demonstrated ability to improve uptime, reliability, and operational processes. ‱ Deep understanding of CI/CD, dev exp, infrastructure-as-code, and automation. ‱ Experience designing on-call processes and incident response frameworks. ‱ Experience managing at least one team of SRE engineers. ‱ Strong communication skills, with the ability to influence across teams. ‱ Experience supporting data platforms and ML systems in production environments. ‱ MLOps experience (model deployment, monitoring, retraining workflows).

đŸ–ïž Benefits

‱ Lead by example, move fast, make data-aware decisions ‱ Continuous push for more, always with a focus on delivering real value to customers

Apply Now

Similar Jobs

🕒 May 20

Replit

51 - 200

đŸ€– Artificial Intelligence

đŸ€ B2B

Join Replit as a Staff Site Reliability Engineer, enhancing performance and reliability of our infrastructure. Collaborate to ensure scalable solutions while mentoring engineers.

đŸ‡ȘđŸ‡ș Europe – Remote

⏰ Full Time

🔮 Lead

⛑ DevOps & Site Reliability Engineer (SRE)

🕒 February 17

Thrill

11 - 50

🎼 Gaming

đŸ„œ AR/VR

Infrastructure/DevOps Engineer responsible for managing AWS and Kubernetes at Thrill Labs. Working on high-scalability projects and improving security measures in a fast-growing tech startup.

đŸ‡ȘđŸ‡ș Europe – Remote

⏰ Full Time

🟠 Senior

🔮 Lead

⛑ DevOps & Site Reliability Engineer (SRE)