Senior Site Reliability Engineer

🕒 July 20

🇬🇧 United Kingdom – Remote

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

👻 Ghost score 41%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Runware

Runware

11 - 50 employees

Founded 2023

🤖 Artificial Intelligence

🔌 API

📱 Media

Artificial Intelligence • API • Media

Runware is a flexible generative AI platform that specializes in high-quality media creation for images and videos through an affordable and fast API. It is capable of advanced tasks such as image generation, video inference, upscaling, and background removal, making it an essential tool for developers seeking to enhance their projects with AI technology. Powered by a custom Sonic Inference Engine™ and renewable energy, Runware ensures quick and efficient media generation without requiring complex infrastructure or machine learning expertise.

📋 Description

• Own and improve the reliability, availability and performance of critical production services across the Runware platform • Define and evolve our reliability practices, including SLIs, SLOs, alerting, observability and production-readiness standards • Investigate complex production issues across distributed systems, APIs, networking, queues, databases and GPU-backed workloads, participating in our engineering on-call rotation • Lead and contribute to incident reviews and RCAs, turning recurring failure modes into lasting engineering improvements • Reduce operational toil through automation, automated remediation and improvements to deployment safety, recovery and system resilience • Work closely with Engineering and DevOps teams on capacity planning, performance, scaling and architectural improvements as the platform grows

🎯 Requirements

• Have strong experience operating and troubleshooting production systems at scale in an SRE, Production Engineering, Platform Engineering or similar role • Have a strong understanding of distributed systems and are comfortable debugging across applications, databases, queues, containers, networking and infrastructure • Have experience designing and operating observability systems using metrics, logs and distributed tracing • Understand SRE principles including SLIs, SLOs, error budgets, capacity planning, incident management and reducing operational toil • Have experience with Kubernetes, containers, IaC and automated deployment practices, alongside the ability to write software and automation using languages such as Python, Go or PHP • Take strong ownership of production problems and are comfortable participating in an engineering on-call rotation, taking issues from initial investigation through to long-term remediation • Bonus • Experience operating high-throughput or low-latency APIs and distributed systems • Experience with bare-metal infrastructure, GPU environments or AI and ML workloads • Experience with RabbitMQ or other distributed messaging and queueing systems • Experience operating MySQL, Redis, ClickHouse or similar production data systems • Experience with global traffic management, load balancing, CDN platforms and hybrid infrastructure environments • Experience building automated scaling, capacity management or self-healing systems

🏖️ Benefits

• Generous paid time off – vacation, sick days, public holidays • Meaningful stock options – share in the upside you create • Remote-first setup – work from home anywhere we can employ you • Flexible hours – own your schedule outside core collaboration blocks • Family leave – paid maternity, paternity, and caregiver time • Company retreats – twice-yearly gatherings in inspiring locations

Apply Now

Similar Jobs

🕒 July 20

Ensono

1001 - 5000

💼 Consulting

Senior DevOps Consultant at Ensono delivering complex projects with deep engineering skills and a focus on quality. Engage in project lifecycle and collaborate with client teams in a remote setting.

🕒 July 14

Ensono

1001 - 5000

💼 Consulting

Senior DevOps Consultant managing complex projects. Overseeing end-to-end ownership of deliverables while collaborating with clients and internal teams.

🕒 July 13

ClickHouse

51 - 200

☁️ SaaS

🏢 Enterprise

🤖 Artificial Intelligence

Site Reliability Engineer at ClickHouse designing systems for real-time analytics. Leading initiatives for cloud infrastructure reliability, availability, and performance.

🕒 July 13

Brahma

11 - 50

₿ Crypto

💳 Fintech

🔌 API

Lead Infrastructure / DevOps Engineer overseeing the AI Platform and Infrastructure team at BRAHMA AI. Guiding a team of engineers in managing high-performance GPU infrastructure and multi-cloud setups.

🕒 July 9

Ensono

1001 - 5000

💼 Consulting

Site Reliability Engineer managing Cloud and Infrastructure as Code at Ensono. Leading client-facing discussions and driving service improvement initiatives.