Senior Site Reliability Engineer, Cloud Platform

Job not on LinkedIn

🔥 0 minutes ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Salve.Inno

Salve.Inno

11 - 50 employees

Founded 2024

💼 Consulting

📣 Marketing

📦 Logistics

Consulting • Marketing • Logistics

Salve. Inno is a recruitment and consulting firm that connects exceptional talent with businesses through personalized hiring strategies and global remote sourcing. The company specializes in recruitment for roles across sectors such as marketing, forex, and iGaming, offering candidate sourcing, screening, and career-site driven hiring experiences while emphasizing DE&I, communication, and innovative process building. Founded in 2024 and headquartered in Gdańsk, Poland, Salve. Inno operates with a small team and a global footprint via remote job listings and consulting services.

📋 Description

• Maintain the reliability, availability, and performance of production and pre-production environments • Monitor platform health and improve alerting, automation, and operational processes • Respond to production incidents, participate in root cause analysis, and implement long-term improvements • Design, build, and enhance observability solutions using metrics, logs, traces, and dashboards • Partner with software engineers to improve application reliability throughout the development lifecycle • Develop and maintain operational documentation, troubleshooting guides, and runbooks • Automate repetitive operational tasks to improve efficiency and reduce manual intervention • Participate in on-call rotations while continuously improving incident response processes • Promote reliability engineering principles, operational excellence, and continuous improvement across engineering teams

🎯 Requirements

• Bachelor's or Master's degree in Engineering, Computer Science, or a related field • Strong experience operating Kubernetes or other container orchestration platforms • Experience supporting large-scale production services • Hands-on experience with AWS • Experience with Prometheus, Grafana, and ELK • Strong scripting skills in Bash, Python, or Go • Experience administering Linux-based production environments • Experience with Infrastructure as Code or configuration management tools such as Terraform or Ansible • Solid understanding of networking fundamentals, including TCP/IP, DNS, load balancing, and routing • Excellent troubleshooting, communication, and collaboration skills • A proactive mindset with a passion for automation and reliability • Nice to have: experience with SIP or VoIP technologies • Nice to have: familiarity with MySQL or PostgreSQL • Nice to have: experience with Redis or other NoSQL databases

🏖️ Benefits

• Long-term, full-time collaboration • Flexible remote working environment • Professional development opportunities, including training and technical learning • Opportunity to work on innovative cloud technologies used by customers worldwide • Collaborative engineering culture focused on knowledge sharing and continuous improvement • Modern Apple equipment provided • Inclusive, respectful workplace

Apply Now

Similar Jobs

🕒 3 days ago

CrowdStrike

5001 - 10000

🔒 Cybersecurity

☁️ SaaS

🤖 Artificial Intelligence

Site Reliability Engineer maintaining CrowdStrike’s large-scale cybersecurity cloud platform. Automating operations, monitoring distributed systems, and leading incident response for reliable 24x7 service.

🕒 August 5

NICE

5001 - 10000

☁️ SaaS

🤖 Artificial Intelligence

📡 Telecommunications

Forward Deployed Engineer building AI agents and full-stack automation for NICE’s customer-experience software. Integrating conversational systems with enterprise platforms and deploying customer self-service solutions.

🕒 August 4

Pliant

201 - 500

💳 Fintech

☁️ SaaS

🤝 B2B

Engineering Manager building Pliant’s Site Reliability function for its B2B payments platform. Establishing SLOs, incident processes, observability, and a new reliability engineering team.

🕒 July 28

PatSnap

501 - 1000

💼 Consulting

🏥 Healthcare

📦 Logistics

Site Reliability Engineering Leader at PatSnap, leading the SRE team ensuring reliability for a global SaaS platform. Overseeing strategy, automation, and team development in cloud technologies.

🕒 July 28

TwinStream

51 - 200

🎖️ Defense

💼 Consulting

📦 Logistics

DevOps Engineer maintaining and deploying cross-domain systems using Docker and AMQP architecture for TwinStream clients. Collaborating with teams and ensuring system performance and availability.