Search Remote Jobs

Staff Site Reliability Engineer

đŸ”„ 10 minutes ago

đŸ‡ȘđŸ‡ș Europe – Remote

⏰ Full Time

🔮 Lead

⛑ DevOps & Site Reliability Engineer (SRE)

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Wand AI

Wand AI

51 - 200 employees

Founded 2022

đŸ€– Artificial Intelligence

🏱 Enterprise

☁ SaaS

Artificial Intelligence ‱ Enterprise ‱ SaaS

Wand AI is an enterprise software company (Wand Synthesis AI Inc. ) that builds the “Agentic Labor Infrastructure” to enable governments and large enterprises to create, manage, and scale hybrid workforces composed of humans and autonomous AI agents. Their platform (branded Wand OS / Agentic Workforce Technology) provides management, oversight, interoperability across systems, security options (SOC2-ready, on‑premise/private cloud/hosted), dashboards, decision tracking, and tools for deploying and governing agentic workflows at scale. Wand positions itself as a B2B/enterprise provider that turns AI into operational labor for regulated and large-scale organizations.

📋 Description

‱ Architect, deploy, and operate scalable, secure production environments (AWS preferred). ‱ Lead reliability improvements across multiple engineering streams. ‱ Design and evolve Kubernetes-based infrastructure, including migration and optimisation initiatives. ‱ Build and enforce strong Infrastructure-as-Code standards. ‱ Define and operationalise SLIs, SLOs, and error budgets. ‱ Strengthen observability across applications, infrastructure, data pipelines, and ML systems. ‱ Work closely with product and data teams to integrate model analytics and product telemetry into reliability insights. ‱ Work across and optimise the entire CI/CD pipeline, from build to deploy to rollback. ‱ Improve release safety, deployment frequency, and predictability of SLAs. ‱ Lead incident response for complex cross-system failures and drive postmortems. ‱ Reduce operational toil through automation and platform engineering improvements. ‱ Design processes and tooling to absorb, standardise, and troubleshoot customer environments. ‱ Support and productionise ML workloads (MLOps practices including model deployment, monitoring, retraining workflows). ‱ Ensure infrastructure aligns with enterprise-grade security and regulatory requirements. ‱ Mentor engineers and raise the overall reliability bar across teams.

🎯 Requirements

‱ Extensive hands-on experience in SRE or Production Engineering roles. ‱ Demonstrated experience building or scaling SRE practices in high-growth or complex environments. ‱ Deep expertise in AWS or Azure-based cloud infrastructure. ‱ Strong experience with Kubernetes (including migration, scaling, and production hardening). ‱ Advanced Infrastructure-as-Code experience (Terraform or equivalent). ‱ End-to-end CI/CD pipeline design and optimisation experience. ‱ Strong experience with observability tooling across distributed systems. ‱ Experience troubleshooting complex multi-tenant or customer-hosted environments. ‱ Experience supporting production data platforms and ML systems. ‱ MLOps experience, including model deployment and monitoring. ‱ Strong understanding of distributed systems, scalability, and fault tolerance. ‱ Systems thinker who understands interactions across infrastructure, product, data, and ML. ‱ Excellent communication skills and ability to work cross-functionally.

đŸ–ïž Benefits

‱ Health insurance ‱ Paid time off ‱ Professional development

Apply Now

Similar Jobs

🕒 May 20

Replit

51 - 200

đŸ€– Artificial Intelligence

đŸ€ B2B

Join Replit as a Staff Site Reliability Engineer, enhancing performance and reliability of our infrastructure. Collaborate to ensure scalable solutions while mentoring engineers.

đŸ‡ȘđŸ‡ș Europe – Remote

⏰ Full Time

🔮 Lead

⛑ DevOps & Site Reliability Engineer (SRE)

🕒 February 17

Thrill

11 - 50

🎼 Gaming

đŸ„œ AR/VR

Infrastructure/DevOps Engineer responsible for managing AWS and Kubernetes at Thrill Labs. Working on high-scalability projects and improving security measures in a fast-growing tech startup.

đŸ‡ȘđŸ‡ș Europe – Remote

⏰ Full Time

🟠 Senior

🔮 Lead

⛑ DevOps & Site Reliability Engineer (SRE)