Senior Site Reliability Engineer – Hiring Globally

Stelle nicht auf LinkedIn

🕒 vor 25 Tagen

🇬🇧 Vereinigtes Königreich – Remote

⏰ Vollzeit

🟠 Senior

⛑ DevOps- und Site Reliability Engineer (SRE)

👻 Geisterscore 15%

infoinfo

🗣️🇺🇸🇬🇧 Englisch erforderlich

Jetzt Bewerben
Ähnliche Remote-Jobs finden

📊 Überprüfen Sie Ihre Lebenslauf-Bewertung für diese Stelle

Verbessern Sie Ihre Chancen auf ein Vorstellungsgespräch, indem Sie Ihre Lebenslauf-Bewertung vor der Bewerbung überprüfen.

Logo of Cognativ

Cognativ

11 - 50 Mitarbeiter

Gegründet 2018

💼 Beratung

🥽 AR/VR

🤖 Künstliche Intelligenz

Consulting • AR/VR • Artificial Intelligence

Cognativ ist eine App-Entwicklungsagentur mit Sitz in Melbourne, Australien. Sie kombinieren Benutzererfahrungsdesign, Produktstrategie und Softwaretechnik, um mobile und Webanwendungen, E-Commerce-Systeme und markengebundene digitale Erlebnisse zu liefern. Ihre Spezialgebiete umfassen UX- und UI-Design, Website- und App-Entwicklung (iOS, Android, React Native, React. js, node. js, Python), Conversion-Optimierung und aufkommende Technologien wie KI/Maschinelles Lernen und AR/VR; sie positionieren sich als qualitätsorientierte IT-Dienstleistungs- und Beratungspartner.

Beschreibung

• Own service objectives. Define SLIs and SLOs for the services that matter, manage error budgets, and use them to drive prioritization and change-rate decisions. Make reliability a measured number, not a feeling. • Make observability trustworthy. Own alert quality end-to-end: high signal-to-noise, alarms that reliably catch the incidents they are meant to catch, and the discipline to turn off a misleading alarm until it is fixed properly. Build the dashboards and the custom metrics, exporters, and instrumentation (CloudWatch, OpenTelemetry) needed to see the system clearly. • Lead incident response. Run incidents calmly, drive mean-time-to-recovery down, and produce blameless postmortems with action items that actually get closed. Improve and own the on-call rotation and its health. • Plan capacity and performance. Forecast and right-size compute (especially GPU), Kafka/MSK throughput and partitioning, RDS/TimescaleDB load, and Redis. Catch saturation before customers do. • Own business continuity and disaster recovery. Backups, replication, failover, and recovery for RDS, MSK, Redis, and the edge fleet. Define RPO/RTO and prove them with regular, tested game days, not assumptions. • Keep the edge fleet healthy. Remote diagnosis and recovery over AWS IoT, container auto-update over systemd timers, and the ongoing CentOS 7 migration to the containerized media stack (Ubuntu 22.04). • Engineer away toil. Write real software (Python, Golang, Bash) to automate operational work, self-heal common failures, and make reliability repeatable instead of heroic. • Govern production change safely. Enforce collaborative, reviewed change management; protect the system from risky, unilateral changes (topology, instance-count, scaling, and config).

🎯 Anforderungen

• AWS certification is mandatory. A current AWS Certified DevOps Engineer – Professional or AWS Certified Solutions Architect – Professional is strongly preferred. • 10+ years in Site Reliability Engineering or production operations at scale. We do not expect mastery of every area below on day one. We expect real depth in several and the ability to ramp quickly on the rest. • Demonstrated SLO/error-budget practice. You have defined SLIs and SLOs, run against an error budget, and used it to make real decisions. • Strong production observability skills. Deep with metrics, logs, alarming, and dashboards (CloudWatch, OpenTelemetry, and/or Prometheus/Grafana/Datadog), and able to build the instrumentation when it does not exist. • Proven incident command. You have led incidents, owned an on-call rotation, and written postmortems that changed how a system behaved. • Capacity planning and performance experience across compute, databases, and a messaging or streaming system (Kafka/MSK ideal). • Disaster recovery ownership: backups, replication, failover, and tested RPO/RTO. • Software engineering ability for automation. Comfortable writing Python, Golang, and Bash to build reliability tooling, not just configure off-the-shelf tools. • Expert with Terraform (or equivalent IaC) and strong Linux administration (shell plus Linux GNU utils), comfortable from cloud to bare-metal/edge. • Database operations experience with PostgreSQL (time-series a plus). • A reliability mindset: you instrument before you guess, and you write the runbook. • Nice to have • Operating GPU workloads and serving computer-vision or ML models in production (CUDA, Deep Learning AMIs, inference scaling). • Apache MSK / Kafka and streaming-data operations (Kinesis, Kinesis Video Streams). • AWS IoT Core at scale: device provisioning, certificates, secure tunnelling. • Managing a fleet of edge / on-premise devices (golden images, remote update, systemd). • Operating and modernizing legacy systems (Java 8, Jetty, CentOS). • Chaos engineering / game-day practice, and capacity modeling. • Familiarity with Bazel in a monorepo; Cloudflare, Cognito/Auth0, API Gateway.

Jetzt Bewerben

Ähnliche Jobs

🕒 vor 25 Tagen

Peratera

11 - 50

💳 Fintech

🤝 B2B

🔌 API

DevOps/SRE Engineer responsible for platform reliability and automation at UK fintech. Building and evolving cloud infrastructure with automation and observability practices.

🇬🇧 Vereinigtes Königreich – Remote

⏰ Vollzeit

🟡 Mittelstufe

🟠 Senior

⛑ DevOps- und Site Reliability Engineer (SRE)

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 1 Monat

GitLab

1001 - 5000

💼 Beratung

📣 Marketing

🤖 Künstliche Intelligenz

Senior Backend Engineer developing cloud-native and self-managed deployment environments for GitLab. Building consistency in deployment across development and production environments.

🇬🇧 Vereinigtes Königreich – Remote

💰 Secondary Market im 2020-11

⏰ Vollzeit

🟠 Senior

⛑ DevOps- und Site Reliability Engineer (SRE)

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 1 Monat

Runware

11 - 50

🤖 Künstliche Intelligenz

🔌 API

📱 Medien

Site Reliability Engineer ensuring reliability and performance of Runware's AI platforms. Collaborating across software, infrastructure, and operations to enhance observability and reduce incidents.

🇬🇧 Vereinigtes Königreich – Remote

⏰ Vollzeit

🟠 Senior

⛑ DevOps- und Site Reliability Engineer (SRE)

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 1 Monat

Ensono

1001 - 5000

💼 Beratung

Senior DevOps Consultant at Ensono delivering complex projects with deep engineering skills and a focus on quality. Engage in project lifecycle and collaborate with client teams in a remote setting.

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 1 Monat

Ensono

1001 - 5000

💼 Beratung

Senior DevOps Consultant managing complex projects. Overseeing end-to-end ownership of deliverables while collaborating with clients and internal teams.

🗣️🇺🇸🇬🇧 Englisch erforderlich