Senior Site Reliability Engineer

Vaga não está no LinkedIn

🕒 Julho 27

🗣️🇺🇸🇬🇧 Inglês obrigatório

Candidatar-se
Encontrar Vagas Remotas Similares

📊 Verifique sua pontuação de currículo para esta vaga

Melhore suas chances de conseguir uma entrevista verificando sua pontuação de currículo antes de se candidatar.

Logo of Cognativ

Cognativ

11 - 50 funcionários

Fundada em 2018

💼 Consultoria

🥽 AR/VR

🤖 Inteligência Artificial

Consulting • AR/VR • Artificial Intelligence

Cognativ é uma agência de desenvolvimento de aplicativos com sede em Melbourne, Austrália. Eles combinam design de experiência do usuário, estratégia de produto e engenharia de software para entregar aplicações móveis e web, sistemas de e-commerce e experiências digitais com marca. Suas especialidades incluem design UX e UI, desenvolvimento de sites e aplicativos (iOS, Android, React Native, React. js, node. js, Python), otimização de conversão e tecnologias emergentes como IA/aprendizado de máquina e AR/VR; posicionam-se como parceiros de serviços de TI e consultoria focados em qualidade.

Descrição

• **What you'll keep reliableYou will own the operational health of all of the following:** • - Computer-vision / AI models: frame-based inference services running on GPU EC2 (g4dn-class, AWS Deep Learning AMIs), fed by camera frames from S3 and a Kafka (Amazon MSK) event bus, with Redis (ElastiCache) for state. Outputs flow through an alerts pipeline (OutgoingInferenceMessage to SNS/IoT to notification workers). • - Python services: the AI/alerts inference tier and supporting tooling. • - Legacy Java services: ~140 Java 8 services and libraries (REST APIs, SQS/SNS workers, Lambda functions) running on Jetty 9.4, deployed to Elastic Beanstalk, ECS, and Lambda. • - Edge appliances ("media boxes"): Ubuntu 22.04 / Docker Compose appliances managed remotely over AWS IoT Core secure tunnelling, with Cloudflare tunnels for egress. Includes an active CentOS 7 to containerized-stack migration effort. • - Data and messaging backbone: PostgreSQL (RDS) across many schemas, TimescaleDB for analytics, Redis, DynamoDB, Amazon MSK (Kafka), SQS/SNS, and Kinesis. • **What you'll do Reliability and operations (the core of the role)** • - Own service objectives. Define SLIs and SLOs for the services that matter, manage error budgets, and use them to drive prioritization and change-rate decisions. Make reliability a measured number, not a feeling. • - Make observability trustworthy. Own alert quality end-to-end: high signal-to-noise, alarms that reliably catch the incidents they are meant to catch, and the discipline to turn off a misleading alarm until it is fixed properly. Build the dashboards and the custom metrics, exporters, and instrumentation (CloudWatch, OpenTelemetry) needed to see the system clearly. • - Lead incident response. Run incidents calmly, drive mean-time-to-recovery down, and produce blameless postmortems with action items that actually get closed. Improve and own the on-call rotation and its health. • - Plan capacity and performance. Forecast and right-size compute (especially GPU), Kafka/MSK throughput and partitioning, RDS/TimescaleDB load, and Redis. Catch saturation before customers do. • - Own business continuity and disaster recovery. Backups, replication, failover, and recovery for RDS, MSK, Redis, and the edge fleet. Define RPO/RTO and prove them with regular, tested game days, not assumptions. • - Keep the edge fleet healthy. Remote diagnosis and recovery over AWS IoT, container auto-update over systemd timers, and the ongoing CentOS 7 migration to the containerized media stack (Ubuntu 22.04). • - Engineer away toil. Write real software (Python, Golang, Bash) to automate operational work, self-heal common failures, and make reliability repeatable instead of heroic. • - Govern production change safely. Enforce collaborative, reviewed change management; protect the system from risky, unilateral changes (topology, instance-count, scaling, and config). • **Delivery and platform (in support of reliability)** • - Keep CI/CD healthy and safe: CircleCI with Bazel/Gradle builds, OIDC-based AWS auth, container builds to ECR, and EB/ECS/Lambda deploys, tuned so that releases are safe, observable, and reversible. • - Maintain Terraform for the AWS estate (compute, networking, IAM, databases, messaging, monitoring), following our module-based conventions and S3-backed state. • - Harden security and compliance: IAM least-privilege, Secrets Manager/KMS, TLS and certificate management (including IoT device certificates), CloudTrail, and AWS Config.

🎯 Requisitos

• **- AWS certification is mandatory. A current AWS Certified DevOps Engineer – Professional or AWS Certified Solutions Architect – Professional is strongly preferred. • - 10+ years in Site Reliability Engineering or production operations at scale. We do not expect mastery of every area below on day one. We expect real depth in several and the ability to ramp quickly on the rest. • - Demonstrated SLO/error-budget practice. You have defined SLIs and SLOs, run against an error budget, and used it to make real decisions. • - Strong production observability skills. Deep with metrics, logs, alarming, and dashboards (CloudWatch, OpenTelemetry, and/or Prometheus/Grafana/Datadog), and able to build the instrumentation when it does not exist. • - Proven incident command. You have led incidents, owned an on-call rotation, and written postmortems that changed how a system behaved. • - Capacity planning and performance experience across compute, databases, and a messaging or streaming system (Kafka/MSK ideal). • - Disaster recovery ownership: backups, replication, failover, and tested RPO/RTO. • - Software engineering ability for automation. Comfortable writing Python, Golang, and Bash to build reliability tooling, not just configure off-the-shelf tools. • - Expert with Terraform (or equivalent IaC) and strong Linux administration (shell plus Linux GNU utils), comfortable from cloud to bare-metal/edge. • - Database operations experience with PostgreSQL (time-series a plus). • - A reliability mindset: you instrument before you guess, and you write the runbook. • Nice to have • - Operating GPU workloads and serving computer-vision or ML models in production (CUDA, Deep Learning AMIs, inference scaling). • - Apache MSK / Kafka and streaming-data operations (Kinesis, Kinesis Video Streams). • - AWS IoT Core at scale: device provisioning, certificates, secure tunnelling. • - Managing a fleet of edge / on-premise devices (golden images, remote update, systemd). • - Operating and modernizing legacy systems (Java 8, Jetty, CentOS). • - Chaos engineering / game-day practice, and capacity modeling. • - Familiarity with Bazel in a monorepo; Cloudflare, Cognito/Auth0, API Gateway.

Candidatar-se

Vagas Similares

🕒 Julho 24

LeadVenture™

1001 - 5000

💼 Consultoria

📣 Marketing

📦 Logística

Site Reliability Engineer focused on monitoring and optimizing applications for LeadVenture, a SaaS provider. Collaborating across teams to ensure seamless user experience with production systems.

🇲🇽 México – Remoto

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Julho 23

Genesys

5001 - 10000

💼 Consultoria

🏥 Saúde

🛡️ Seguros

Drive reliability of enterprise identity services ensuring security and seamless access for global workforce at Genesys. Collaborate with cross-functional teams to enhance IAM reliability and system resilience.

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Julho 23

Sezzle

201 - 500

💳 Fintech

👥 B2C

🛍️ Comércio Eletrônico

Senior Site Reliability Engineer building and upgrading infrastructure solutions for Sezzle. Leading efforts towards enhanced reliability and scalability in a fast-paced environment.

🇲🇽 México – Remoto

💵 $7.000 - $12.000 / mês

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Julho 20

Truelogic Software

501 - 1000

☁️ SaaS

🤝 B2B

🏢 Corporativo

Semi-senior DevOps Engineer involved in designing and maintaining AWS infrastructure for a SaaS company. Collaborating with cross-functional teams to enhance platform reliability and performance.

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Julho 16

3Pillar Global

1001 - 5000

💼 Consultoria

🏥 Saúde

🛡️ Seguros

Senior DevOps Engineer at 3Pillar, ensuring seamless integration and deployment of AI-native products. Leading a team in optimizing deployment strategies and implementing cutting-edge technologies.

🇲🇽 México – Remoto

💰 Private Equity Round em 2021-10

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório