Senior Site Reliability Engineer

🕒 vor 1 Monat

🇬🇧 Vereinigtes Königreich – Remote

⏰ Vollzeit

🟠 Senior

⛑ DevOps- und Site Reliability Engineer (SRE)

👻 Geisterscore 30%

infoinfo

🗣️🇺🇸🇬🇧 Englisch erforderlich

Jetzt Bewerben
Ähnliche Remote-Jobs finden

📊 Überprüfen Sie Ihre Lebenslauf-Bewertung für diese Stelle

Verbessern Sie Ihre Chancen auf ein Vorstellungsgespräch, indem Sie Ihre Lebenslauf-Bewertung vor der Bewerbung überprüfen.

Logo of ClickHouse

ClickHouse

51 - 200 Mitarbeiter

Gegründet 2016

☁️ SaaS

🏢 Unternehmen

🤖 Künstliche Intelligenz

SaaS • Enterprise • Artificial Intelligence

ClickHouse ist ein schnelles, ressourceneffizientes Echtzeit‑Data‑Warehouse und eine Open‑Source‑Datenbank, die für überragende Abfrageleistung in geschäfts‑ und zeitkritischen Anwendungen entwickelt wurde. Es ist als Cloud‑Service auf führenden Plattformen wie AWS, GCP und Azure verfügbar, bietet eine Option „Bring Your Own Cloud“ und eine breite Palette an Integrationen für den nahtlosen Betrieb in unterschiedlichen Tech‑Stacks. ClickHouse überzeugt bei Echtzeit‑Analysen, Machine Learning, Business Intelligence und Observability und ist damit eine ideale Wahl für Anwendungsfälle wie Financial Services, Fraud Detection und Gaming‑Analytics. Es unterstützt entwicklerfreundliche SQL‑Operationen, bietet kosteneffiziente Storage‑Lösungen und stellt eine Open‑Source‑Alternative zu traditionellen Datenbanken dar. Unternehmen wie Sony, Lyft, Cisco, GitLab und Twilio setzen ClickHouse wegen seiner Skalierbarkeit, Effizienz und Benutzerfreundlichkeit ein.

Beschreibung

• Collaborate with various engineering teams in ClickHouse to design and implement scalable, secure, and highly available systems for ClickHouse. • Establish and manage service level objectives (SLOs) and service level agreements (SLAs) for ClickHouse Cloud. • Ensure all the infrastructure components in ClickHouse Cloud (including Dataplane, Control Plane and ClickHouse Core) have monitoring and alerting in place to ensure timely detection and resolution of incidents. • Enhance and refine incident response processes and post-mortem analysis for any outages in ClickHouse Cloud including working with the support team to communicate to the impacted customers. • Continuously improve the reliability and performance of our ClickHouse services. • Plan, enable, and drive Chaos initiatives across Engineering teams, based upon internal priorities. • Manage on-call processes to respond to performance and reliability issues, and establish best practices for coordinating escalation to resolve issues and minimize downtime.

🎯 Anforderungen

• Bachelor’s or Master’s degree in Computer Science or a related field. • At least 8 years of experience in Site Reliability Engineering or a related field. • Previous experience using ClickHouse in production. • Hands on experience with Go and/or Python. • Strong knowledge of cloud computing platforms such as AWS, Azure, or Google Cloud Platform. • Excellent understanding of distributed databases and SQL, particularly ClickHouse is a major plus. • Hands on experience with container orchestration tools such as Kubernetes or Docker Swarm. • Strong experience with automation and configuration management tools such as Ansible, Terraform, or Puppet. • You are a strong problem solver and have solid production debugging skills. • You are passionate about efficiency, availability, scalability, and data governance. • You thrive in a fast paced environment, and see yourself as a partner with the business with the shared goal of moving the business forward. • You have a high level of responsibility, ownership, and accountability. • Excellent communication and interpersonal skills.

🏖️ Vorteile

• Flexible work environment - ClickHouse is a globally distributed company and remote-friendly. We currently operate in over 20 countries. • Healthcare - Employer contributions towards your healthcare. • Equity in the company - Every new team member who joins our company receives stock options. • Time off - Flexible time off in the US, generous entitlement in other countries. • A $500 Home office setup if you’re a remote employee. • Global Gatherings – We believe in the power of in-person connection and offer opportunities to engage with colleagues at company-wide offsites.

Jetzt Bewerben

Ähnliche Jobs

🕒 vor 1 Monat

Brahma

11 - 50

₿ Crypto

💳 Fintech

🔌 API

Lead Infrastructure / DevOps Engineer overseeing the AI Platform and Infrastructure team at BRAHMA AI. Guiding a team of engineers in managing high-performance GPU infrastructure and multi-cloud setups.

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 1 Monat

Ensono

1001 - 5000

💼 Beratung

Site Reliability Engineer managing Cloud and Infrastructure as Code at Ensono. Leading client-facing discussions and driving service improvement initiatives.

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 1 Monat

Omilia - Conversational Intelligence

201 - 500

💼 Beratung

🛡️ Versicherung

✈️ Reisen

Senior Site Reliability Engineer operating and maintaining production clusters while developing observability solutions. Collaborating with teams to enhance platform reliability through automation and monitoring.

🇬🇧 Vereinigtes Königreich – Remote

⏰ Vollzeit

🟠 Senior

⛑ DevOps- und Site Reliability Engineer (SRE)

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 1 Monat

DeepHealth

11 - 50

🏥 Gesundheitswesen

💼 Beratung

📦 Logistik

DevOps Engineer managing AWS infrastructure and enhancing platform reliability at DeepHealth. Collaborating with teams to automate processes and improve software delivery.

🇬🇧 Vereinigtes Königreich – Remote

💵 £60.000 - £70.000 / Jahr

💰 €225.000 Grant im 2019-08

⏰ Vollzeit

🟡 Mittelstufe

🟠 Senior

⛑ DevOps- und Site Reliability Engineer (SRE)

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 1 Monat

DeepHealth

11 - 50

🏥 Gesundheitswesen

💼 Beratung

📦 Logistik

DevOps Engineer responsible for AWS and Kubernetes platform management at DeepHealth, a healthcare SaaS provider. Ensuring cloud infrastructure is secure and reliable for efficient software delivery.

🇬🇧 Vereinigtes Königreich – Remote

💵 £60.000 - £70.000 / Jahr

💰 €225.000 Grant im 2019-08

⏰ Vollzeit

🟡 Mittelstufe

🟠 Senior

⛑ DevOps- und Site Reliability Engineer (SRE)

🗣️🇺🇸🇬🇧 Englisch erforderlich