Site Reliability Engineer

🕒 July 3

🇮🇳 India – Remote

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

👻 Ghost score 52%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Moniepoint Inc. (Formerly TeamApt Inc.)

Moniepoint Inc. (Formerly TeamApt Inc.)

1001 - 5000 employees

💳 Fintech

🏦 Banking

Fintech • Banking • Payments

Moniepoint Inc. is Africa's all-in-one financial ecosystem that provides seamless solutions in payments, banking, credit, and business management for over 10 million businesses and individuals. Operating as Nigeria's largest merchant acquirer, Moniepoint powers the majority of Point of Sale (POS) transactions in the country. The company processes $17 billion monthly while ensuring profitable operations. With operations starting in 2019, Moniepoint continues to support businesses through its comprehensive financial services platform, making significant strides in financial inclusion across emerging markets.

📋 Description

• Participate in on-call rotations as the primary technical lead • Act as Incident Commander during major severity incidents, initiating war rooms, coordinating cross-functional teams, and providing status updates • Instrument code to expose high-cardinality metrics and distributed traces • Define, measure, and defend Service Level Objectives (SLOs) and Error Budgets with product owners • Write production-ready Java, Go, or Python code for internal tooling, automation platforms, and self-healing mechanisms • Partner with Product Engineering teams during design to embed reliability, scalability, and observability patterns • Analyze system performance and traffic patterns to model future capacity needs • Conduct load testing and chaos engineering experiments to verify resilience

🎯 Requirements

• 3–6 years of experience in SRE or Backend Engineering • Strong ability to write clean, performant, and tested code in Java, Go, Rust, or Python • Deep understanding of distributed systems architecture and design patterns • Strong command of microservices fundamentals and event-driven architectures • Experience with Google Cloud Platform (GCP) or similar cloud providers such as AWS/Azure • Proficiency running production workloads on Kubernetes (GKE/EKS) • Ability to troubleshoot cluster and infrastructure issues • Experience designing observability strategies using OpenTelemetry, Prometheus, New Relic, Datadog, or SigNoz • Familiarity with operating and tuning PostgreSQL or MySQL in high-throughput environments • Familiarity with Kafka or RabbitMQ in high-throughput environments

🏖️ Benefits

• Learning and development-focused environment • Knowledge sharing, training, and regular internal technical talks • Attractive salary • Pension • Health insurance • Annual bonus • Other benefits • Inclusive and diverse work environment

Apply Now

Similar Jobs

🕒 July 1

Empower

10,000+ employees

💸 Finance

💳 Fintech

👥 B2C

Senior SRE architecting reliable AWS and Kubernetes infrastructure for Empower’s financial services platform. Leading incident response, automation, observability, security, and engineer mentorship.

AWS

Cloud

Consul

Flux

Kubernetes

Python

Splunk

Terraform

Go

🕒 June 8

RevMind srl

11 - 50

🤖 Artificial Intelligence

💊 Pharmaceuticals

🏢 Enterprise

Senior Cloud DevOps Engineer owning secure, scalable AWS infrastructure for Revmind Labs AI's enterprise AI and analytics systems. Automating deployments, observability, security, and production reliability.

Amazon Redshift

AWS

Cloud

Docker

DynamoDB

EC2

ETL

Flask

Java

JavaScript

Microservices

MySQL

Node.js

NoSQL

Python

React

React Native

Scala

Terraform

🕒 June 1

OpenAI

201 - 500

🤖 Artificial Intelligence

☁️ SaaS

🏢 Enterprise

Partner AI Deployment Engineer responsible for AWS deployment strategies and technical leadership in OpenAI. Guiding enterprise customers from ideation to production while influencing joint account strategy.

AWS

🕒 May 21

Akamai Technologies

5001 - 10000

🔒 Cybersecurity

🏢 Enterprise

📱 Media

Senior Site Reliability Engineer focusing on developing solutions for automation and efficiency with Akamai's Compute products. Enhance reliability and operational excellence in customer-facing applications and infrastructure.

Ansible

AWS

Azure

Cloud

Distributed Systems

Google Cloud Platform

Grafana

Prometheus

Python

SaltStack

Splunk

Terraform

Go

🕒 May 18

Akamai Technologies

5001 - 10000

🔒 Cybersecurity

🏢 Enterprise

📱 Media

Site Reliability Engineer diagnosing networking problems and ensuring performance of cloud solutions at Akamai. Collaborating with Engineering teams to enhance service reliability and efficiency.

Ansible

Chef

Distributed Systems

DNS

Docker

Firewalls

Grafana

HAProxy

Java

Kubernetes

Linux

NGINX

Perl

Prometheus

Puppet

Python

SaltStack

SQL

Terraform