Site Reliability Engineer

Job not on LinkedIn

🔥 0 minutes ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of ArangoDB

ArangoDB

51 - 200 employees

Founded 2015

🏥 Healthcare

📦 Logistics

💼 Consulting

💰 $27.8M Series B on 2021-10

Healthcare • Logistics • Consulting

ArangoDB is a flexible and scalable graph database platform designed for a variety of applications, particularly in a data-intensive environment like Generative AI. It supports multiple data models including graph, vector, document, full-text search, and geospatial databases, all within a single unified system. This allows developers to efficiently navigate complex data interactions and perform advanced analytics, making it ideal for sectors such as healthcare, finance, telecommunications, and more.

📋 Description

• Design, implement, and maintain cloud infrastructure on AWS and Google Cloud platforms • Ensure scalability, performance, and reliability of Kubernetes-based distributed database systems • Collaborate with developers to write production-grade Golang code for infrastructure automation and system operations • Optimize and automate CI/CD pipelines, deployment processes, and monitoring systems • Develop disaster recovery, high availability, and fault-tolerance strategies • Identify system bottlenecks and troubleshoot issues across networking, operating systems, and cloud infrastructure • Implement monitoring, logging, and alerting systems for system health and performance visibility • Participate in on-call rotations and respond to production incidents • Collaborate with cross-functional teams to improve system reliability and scalability • Collaborate with the Customer Success team to resolve customer issues

🎯 Requirements

• Proven experience as an SRE or DevOps Engineer in a cloud-native environment • Minimum 3 years of Kubernetes experience in a production environment • Minimum 3 years of experience deploying and managing production-level cloud resources • Proficiency with Kubernetes for large-scale distributed systems • Experience with AWS and Google Cloud (GCP) • Understanding of networking, security practices, troubleshooting methods, and Linux internals • Familiarity with Docker and containerization technologies • Knowledge of CI/CD practices and tools such as Jenkins and CircleCI • Familiarity with monitoring, alerting, and observability tools such as Prometheus, Grafana, and ELK stack • Knowledge of version control systems, particularly Git • Familiarity with Golang or Python; willingness and capability to learn and apply Golang accepted • Ability to self-organize and work independently as part of a remote team • Experience managing distributed databases or large-scale data storage systems is a plus • Knowledge of cloud security best practices is a plus • Experience with Python or Bash scripting is a plus • Terraform/IaC experience is a plus • GitOps experience is a plus • Strong Golang programming skills and automation-tool development experience are a plus

🏖️ Benefits

• Remote work • Collaboration with experienced engineers, marketers, and product leaders • Opportunity to contribute to cutting-edge AI and data infrastructure • Opportunity to help shape how enterprises build AI-powered applications • Inclusive, growth-oriented team environment

Apply Now

Similar Jobs

🕒 2 days ago

Devoteam

5001 - 10000

💼 Consulting

🏥 Healthcare

📣 Marketing

Data AWS DevSecOps professional evolving secure cloud data platforms for Devoteam’s large-organization clients. Advising teams on architecture, security, governance and observability in a remote Spain-based role.

Airflow

AWS

Cloud

DynamoDB

ETL

Hadoop

Python

Spark

SQL

Terraform

🕒 2 days ago

Logicalis Spain

1001 - 5000

💼 Consulting

DevOps Engineer operando y automatizando plataformas Kubernetes para Logicalis Spain, proveedor de servicios IT empresariales. Mejorando CI/CD, infraestructura cloud y servicios gestionados.

🗣️🇪🇸 Spanish Required

Ansible

AWS

Azure

Cloud

ElasticSearch

Google Cloud Platform

Grafana

Jenkins

Kubernetes

OpenShift

Prometheus

Python

Terraform

🕒 2 days ago

Tempo Software

201 - 500

☁️ SaaS

🏢 Enterprise

⚡ Productivity

Senior Site Reliability Engineer building AWS infrastructure, CI/CD pipelines, and Kubernetes platforms for Tempo’s enterprise productivity software. Automating reliability, observability, security, and cloud deployments.

Ansible

AWS

Cloud

Docker

Java

Kotlin

Kubernetes

Linux

Terraform

🕒 August 7

Affirm

1001 - 5000

💳 Fintech

👥 B2C

🛍️ eCommerce

Senior SRE strengthening reliability for Affirm’s honest, flexible buy-now-pay-later platform. Building incident lifecycle, observability and resilient backend practices for global engineering teams.

AWS

Distributed Systems

Kotlin

Kubernetes

MySQL

Python

🕒 August 5

Miratech

501 - 1000

🤝 B2B

💼 Consulting

☁️ SaaS

Senior DevOps Engineer operating production Kubernetes platforms and GitOps workflows for Miratech, a global IT services and consulting company. Improving cloud infrastructure, CI/CD, observability, and production reliability.

AWS

Cloud

Distributed Systems

DNS

Flux

Google Cloud Platform

Grafana

Jenkins

Kubernetes

Linux

Microservices

Prometheus

Python

Terraform

Go