Site Reliability Engineer – SRE

Job not on LinkedIn

🔥 14 hours ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Rocket.net

Rocket.net

11 - 50 employees

Founded 2020

☁️ SaaS

🛍️ eCommerce

🏢 Enterprise

SaaS • eCommerce • Enterprise

Rocket. net is a managed WordPress hosting platform that delivers fast, secure, and fully-optimized hosting powered by Cloudflare Enterprise. The service offers all-in-one features including enterprise CDN, website firewall (WAF), malware protection, automatic updates, free migrations, and performance tools tailored for WordPress, WooCommerce, agencies, resellers, and enterprise customers, backed by 24/7 expert support. Rocket. net emphasizes ease of use, high speed, and strong security to help businesses, publishers, and online stores improve conversions and reliability.

📋 Description

• Monitor the health, availability, and performance of Rocket.net servers, services, and customer environments. • Proactively identify infrastructure issues, performance degradation, and potential service disruptions. • Investigate alerts and operational events to maintain platform stability. • Perform regular platform health checks and ensure critical systems are operating correctly. • Participate in incident response and coordinate troubleshooting during customer-impacting events. • Communicate platform issues, updates, and resolutions to relevant internal teams. • Provide advanced technical support for VIP customers and customers with complex hosting-related issues. • Act as a senior escalation point for WordPress Support Engineers when issues require deeper technical investigation. • Troubleshoot complex issues involving servers, websites, networking, DNS, performance, caching, and hosting infrastructure. • Assist customers with advanced technical problems beyond standard WordPress troubleshooting. • Investigate and resolve issues involving server resources, application performance, connectivity, and platform behavior. • Work directly with customers when required to provide expert-level technical assistance. • Ensure escalated customer issues are handled with urgency, ownership, and clear communication. • Troubleshoot and maintain Linux-based production environments. • Investigate issues related to NGINX, Apache, PHP-FPM, MySQL/MariaDB, Redis, and other platform services. • Assist with server maintenance, configuration changes, and operational improvements. • Support security updates, system hardening, and infrastructure best practices. • Monitor resource usage and identify capacity or performance concerns. • Help improve monitoring, automation, and operational workflows. • Work closely with WordPress Support Engineers, Shift Leads, Site Reliability Engineers, and Engineering teams. • Provide technical guidance and knowledge sharing to Support teams. • Help create internal documentation, troubleshooting guides, and knowledge base articles. • Identify recurring issues and recommend improvements to reduce future incidents. • Participate in incident reviews and root cause analysis.

🎯 Requirements

• 3+ years of experience in SRE, DevOps, Platform Engineering, or similar roles. • Strong experience troubleshooting Linux production environments. • Experience supporting customer-facing technical environments. • Strong understanding of web hosting technologies including NGINX, Apache, PHP-FPM, MySQL/MariaDB, and Redis. • Advanced troubleshooting skills across WordPress, servers, DNS, networking, and performance issues. • Experience with Linux command line (SSH). • Strong understanding of DNS, HTTP/HTTPS, SSL/TLS, CDN, and caching technologies. • Experience with Cloudflare, WAF, and web performance optimization. • Experience with monitoring tools and incident response processes. • Ability to troubleshoot complex issues independently and communicate technical solutions clearly. • Excellent written and verbal communication skills (English). • Ability to work under pressure during customer-impacting incidents.

🏖️ Benefits

• Ability to work from anywhere in the world. • Flexible Vacation. • Paid Education.

Apply Now

Similar Jobs

🔥 14 hours ago

Branch

501 - 1000

💼 Consulting

📣 Marketing

🔌 API

Senior Site Reliability Engineer at Branch improving platform reliability and scalability with automation. Collaborating with developers to enhance performance and observability of services.

Docker

Gradle

Java

Kubernetes

Spring

Spring Boot

SpringBoot

Terraform

Go

🔥 14 hours ago

Branch

501 - 1000

💼 Consulting

📣 Marketing

🔌 API

Senior Database Reliability Engineer managing MySQL and CloudSQL systems. Ensuring high performance and availability in a remote role at Branch.

Cloud

Google Cloud Platform

MySQL

Redis

SQL

🔥 15 hours ago

Neural Earth

11 - 50

🤖 Artificial Intelligence

🛡️ Insurance

🏠 Real Estate

DevOps Lead at Neural Earth responsible for maintaining CI/CD pipelines and monitoring AWS cloud infrastructure. Support incident response and operational work for smooth engineering processes.

AWS

Azure

Cloud

Cyber Security

Docker

Google Cloud Platform

Kubernetes

Python

Terraform

🔥 16 hours ago

PerfectServe

201 - 500

🏥 Healthcare

⚕️ Healthcare Insurance

☁️ SaaS

Forward Deployment Engineer, AI optimizing and deploying AI voice agent solutions for PerfectServe's customers. Collaborating with Product and Customer Success to ensure successful go-live and ongoing support.

🇺🇸 United States – Remote

💵 $120k - $140k / year

💰 Private Equity Round on 2018-05

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

AWS

Cloud

Python

🔥 17 hours ago

Outlook Group

201 - 500

🏭 Manufacturing

📦 Logistics

🍽️ Food & Beverage

Senior DevOps Engineer overseeing AWS infrastructure and DevOps practices at Outlook Amusements. Collaborating with teams to enhance automation and service reliability in production environments.

Apache

AWS

Azure

Cloud

EC2

Linux

MS SQL Server

MySQL

SQL