Site Reliability Engineer – II

🕒 July 13

🇮🇳 India – Remote

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

👻 Ghost score 14%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of MRSOOL | مرسول

MRSOOL | مرسول

201 - 500 employees

Founded 2015

🍽️ Food & Beverage

✈️ Travel

💼 Consulting

Food & Beverage • Travel • Consulting

MRSOOL is one of the largest delivery platforms in the region, offering an on-demand experience with high user ratings in both Apple's App Store and Google's Play store. MRSOOL provides a "order anything from anywhere" service backed by a large fleet of registered couriers. It enables businesses to access a vast user base, facilitating transformation to on-demand eCommerce and offering a flexible bidding system for service prices. MRSOOL also offers a personalized delivery experience with real-time tracking and communication with couriers, making it a leading option for ordering from local shops, groceries, and restaurants directly to your door.

📋 Description

• Collaborate with development teams to design and implement scalable Infrastructure. • Collaborate with development teams to design and implement automated deployment and testing pipelines. • Develop and maintain monitoring and alerting systems to proactively identify and address issues. • Troubleshoot and escalate production incidents to minimize downtime and improve system reliability. • Continuously improve our infrastructure and processes to optimize scalability and efficiency. • Participate and take ownership for on-call rotations as needed to ensure 24/7 support for our application. • Perform routine maintenance and upgrades as needed to keep our systems up to date. • Contribute to ongoing efforts to improve our security posture and compliance with industry standards. • Communicate complex technical concepts clearly and concisely to both technical and non-technical stakeholders in order to make the right decision. • Mentor and coach junior engineers, fostering their professional growth and enabling them to deliver high-quality work. • Stay up-to-date with the latest advancements and trends in site reliability engineering and share knowledge and insights with the team. • Identify opportunities for organizational enhancements and propose alternatives to optimize team structures and execution.

🎯 Requirements

• Bachelor’s degree in Computer Engineering, Computer Science, or related field. • 5+ years of experience in a similar role, preferably with experience in a high-traffic, high-availability environment. • Proficiency in at least one programming language (Python, Ruby, Java, Go, etc.). • Strong understanding of cloud infrastructure and related technologies (AWS, GCP, Azure, Kubernetes, Docker, etc.) • Excellent troubleshooting and problem-solving skills. • Experience with one or more automation and configuration management tools (Chef, Ansible, Puppet, Terraform, etc.). • Familiarity with monitoring and alerting tools (Prometheus, Grafana, Nagios, etc.) • Strong communication and interpersonal skills, enabling effective collaboration with cross-functional teams. • Ability to navigate ambiguity, set clear expectations, and thrive in a fast-paced, dynamic environment. • A strong grasp of computer science fundamentals when it comes to dealing with distributed systems and networks.

🏖️ Benefits

• Inclusive and Diverse Environment: We foster an inclusive and diverse workplace that values innovation and offers remote environments. • Competitive Compensation: Our compensation packages are highly competitive and include potential share options for certain roles. • Personal Growth and Development: We are committed to your personal and professional growth, providing regular training and an annual learning stipend to help you advance your career in a dynamic environment. • Autonomy and Mentorship: You'll enjoy a high degree of autonomy in your role, supported by mentorship and ambitious goals that pave the way for both your success and the company's growth.

Apply Now

Similar Jobs

🕒 July 13

Akamai Technologies

5001 - 10000

🔒 Cybersecurity

Senior Site Reliability Engineer enhancing the reliability and performance of distributed content delivery infrastructure. Collaborating across teams and implementing robust operational processes for cloud technologies.

Cloud

Grafana

JavaScript

Linux

Oracle

Prometheus

Python

SQL

Unix

🕒 July 9

DBSync

51 - 200

💼 Consulting

🏥 Healthcare

📦 Logistics

Forward Deployment Engineer at DBSync solving tasks for Cloud technology users and ensuring customer success through technical expertise.

AWS

Cloud

Java

Python

SOAP

SQL

Tableau

Go

🕒 July 8

MariaDB

201 - 500

🏢 Enterprise

QA and release engineer testing MariaDB MaxScale, a database proxy for MariaDB clusters. Managing releases, Linux packages, CI/CD automation, and build infrastructure.

Cloud

Jenkins

Linux

MariaDB

MySQL

Python

SQL

🕒 July 8

Pythian

201 - 500

💼 Consulting

🏥 Healthcare

📦 Logistics

Site Reliability Engineer at Pythian focusing on operating large-scale distributed systems. Responsible for designing, deploying, and operating infrastructure with strong collaboration across teams.

Cloud

Docker

Grafana

Kubernetes

Linux

Microservices

Prometheus

Python

Shell Scripting

Terraform

Go

🕒 July 7

Resilinc

201 - 500

💼 Consulting

📦 Logistics

🏥 Healthcare

Site Reliability Engineer responsible for platform availability and automation in cloud environments at Resilinc. Focused on leveraging agentic AI for impactful supply chain solutions.

Azure

Cloud

Distributed Systems

DNS

Docker

Grafana

Hadoop

HDFS

Kafka

Kubernetes

Linux

Postgres

Redis