Site Reliability Engineer – SRE

Job not on LinkedIn

🔥 46 minutes ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Rocket.net

Rocket.net

11 - 50 employees

Founded 2020

☁️ SaaS

🛍️ eCommerce

🏢 Enterprise

SaaS • eCommerce • Enterprise

Rocket. net is a managed WordPress hosting platform that delivers fast, secure, and fully-optimized hosting powered by Cloudflare Enterprise. The service offers all-in-one features including enterprise CDN, website firewall (WAF), malware protection, automatic updates, free migrations, and performance tools tailored for WordPress, WooCommerce, agencies, resellers, and enterprise customers, backed by 24/7 expert support. Rocket. net emphasizes ease of use, high speed, and strong security to help businesses, publishers, and online stores improve conversions and reliability.

📋 Description

• Monitor the health, availability, and performance of Rocket.net servers, services, and customer environments. • Proactively identify infrastructure issues, performance degradation, and potential service disruptions. • Investigate alerts and operational events to maintain platform stability. • Perform regular platform health checks and ensure critical systems are operating correctly. • Participate in incident response and coordinate troubleshooting during customer-impacting events. • Communicate platform issues, updates, and resolutions to relevant internal teams. • Provide advanced technical support for VIP customers and customers with complex hosting-related issues. • Act as a senior escalation point for WordPress Support Engineers when issues require deeper technical investigation. • Troubleshoot complex issues involving servers, websites, networking, DNS, performance, caching, and hosting infrastructure. • Assist customers with advanced technical problems beyond standard WordPress troubleshooting. • Investigate and resolve issues involving server resources, application performance, connectivity, and platform behavior. • Work directly with customers when required to provide expert-level technical assistance. • Ensure escalated customer issues are handled with urgency, ownership, and clear communication. • Troubleshoot and maintain Linux-based production environments. • Investigate issues related to NGINX, Apache, PHP-FPM, MySQL/MariaDB, Redis, and other platform services. • Assist with server maintenance, configuration changes, and operational improvements. • Support security updates, system hardening, and infrastructure best practices. • Monitor resource usage and identify capacity or performance concerns. • Help improve monitoring, automation, and operational workflows. • Work closely with WordPress Support Engineers, Shift Leads, Site Reliability Engineers, and Engineering teams. • Provide technical guidance and knowledge sharing to Support teams. • Help create internal documentation, troubleshooting guides, and knowledge base articles. • Identify recurring issues and recommend improvements to reduce future incidents. • Participate in incident reviews and root cause analysis.

🎯 Requirements

• 3+ years of experience in SRE, DevOps, Platform Engineering, or similar roles. • Strong experience troubleshooting Linux production environments. • Experience supporting customer-facing technical environments. • Strong understanding of web hosting technologies including NGINX, Apache, PHP-FPM, MySQL/MariaDB, and Redis. • Advanced troubleshooting skills across WordPress, servers, DNS, networking, and performance issues. • Experience with Linux command line (SSH). • Strong understanding of DNS, HTTP/HTTPS, SSL/TLS, CDN, and caching technologies. • Experience with Cloudflare, WAF, and web performance optimization. • Experience with monitoring tools and incident response processes. • Ability to troubleshoot complex issues independently and communicate technical solutions clearly. • Excellent written and verbal communication skills (English). • Ability to work under pressure during customer-impacting incidents.

🏖️ Benefits

• Ability to work from anywhere in the world. • Flexible Vacation. • Paid Education.

Apply Now

Similar Jobs

🔥 50 minutes ago

Branch

501 - 1000

💼 Consulting

📣 Marketing

🔌 API

Senior Site Reliability Engineer at Branch improving platform reliability and scalability with automation. Collaborating with developers to enhance performance and observability of services.

🔥 50 minutes ago

Branch

501 - 1000

💼 Consulting

📣 Marketing

🔌 API

Senior Database Reliability Engineer managing MySQL and CloudSQL systems. Ensuring high performance and availability in a remote role at Branch.

🔥 2 hours ago

Neural Earth

11 - 50

🤖 Artificial Intelligence

🛡️ Insurance

🏠 Real Estate

DevOps Lead at Neural Earth responsible for maintaining CI/CD pipelines and monitoring AWS cloud infrastructure. Support incident response and operational work for smooth engineering processes.

🔥 2 hours ago

PerfectServe

201 - 500

🏥 Healthcare

⚕️ Healthcare Insurance

☁️ SaaS

Forward Deployment Engineer, AI optimizing and deploying AI voice agent solutions for PerfectServe's customers. Collaborating with Product and Customer Success to ensure successful go-live and ongoing support.

🇺🇸 United States – Remote

💵 $120k - $140k / year

💰 Private Equity Round on 2018-05

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🔥 3 hours ago

Outlook Group

201 - 500

🏭 Manufacturing

📦 Logistics

🍽️ Food & Beverage

Senior DevOps Engineer overseeing AWS infrastructure and DevOps practices at Outlook Amusements. Collaborating with teams to enhance automation and service reliability in production environments.