
51 - 200 employees
Founded 2009
☁️ SaaS
🔐 Security
🌐 Web 3
SaaS • Security • Web 3
CloudLinux is a leading provider of operating systems designed specifically for web hosting environments. The company offers a line of products including CloudLinux OS Legacy, CloudLinux OS Shared Pro, and CloudLinux OS Solo, each tailored to improve server stability, security, and performance. With features such as kernel live patching, advanced automation and monitoring tools, and specialized WordPress optimization, CloudLinux helps hosting companies maximize security and profitability while ensuring stable server environments. Over 4,000 companies trust CloudLinux to power millions of websites worldwide, benefitting from increased stability and reduced churn rates. The company emphasizes compatibility with major hosting control panels and CentOS, offering solutions for shared hosts, agencies, and small businesses.
🔥 0 minutes ago
🌐 Poland, Romania, +4 more countries – Remote
⏰ Full Time
🟠 Senior
⛑ DevOps & Site Reliability Engineer (SRE)
👻 Ghost score 10%
Improve your chances of getting an interview by checking your resume score before you apply.

51 - 200 employees
Founded 2009
☁️ SaaS
🔐 Security
🌐 Web 3
SaaS • Security • Web 3
CloudLinux is a leading provider of operating systems designed specifically for web hosting environments. The company offers a line of products including CloudLinux OS Legacy, CloudLinux OS Shared Pro, and CloudLinux OS Solo, each tailored to improve server stability, security, and performance. With features such as kernel live patching, advanced automation and monitoring tools, and specialized WordPress optimization, CloudLinux helps hosting companies maximize security and profitability while ensuring stable server environments. Over 4,000 companies trust CloudLinux to power millions of websites worldwide, benefitting from increased stability and reduced churn rates. The company emphasizes compatibility with major hosting control panels and CentOS, offering solutions for shared hosts, agencies, and small businesses.
• Define what “working” means for approximately 70 components • Run SLI definition with squad leads and senior engineers and facilitate ownership sign-off • Build service, fleet, control-efficacy, delivery, and pipeline SLI taxonomies • Attach an SLO, error budget, and owning squad to each SLI, with appropriate tiering • Design and build a push-based, sampled, privacy-constrained telemetry collection pipeline with a defended cardinality budget • Extend agent-side and service-side instrumentation in Python, Go, and Rust • Consolidate dashboards, ad-hoc queries, and reporting paths into a defensible instrument set • Build symptom-based, SLO-anchored alerting with multi-window burn-rate semantics • Establish page, ticket, and dashboard alert tiers and define paging criteria • Ensure every alert has an owner, runbook, and documented failure mode • Maintain alert hygiene through quarterly reviews, deletion, and actionable-rate tracking • Maintain a machine-readable component-to-squad ownership map wired into alert routing • Design severity matrices, acknowledgement SLAs, follow-the-sun rotas across UTC−5 to UTC+8, and handoff protocols • Practice incident command and blameless postmortems within 24 hours • Design escalation so squads carry their own pagers while you operate the platform and coach teams • Deliver first-year outcomes including component inventory, pilot instrumentation, production collection pipeline, tier-1 escalation, full component SLI coverage, alert metrics, squad on-call, and mean time to detect silent control degradation under 24 hours
• Substantial production-engineering or SRE experience, including at least one environment where you defined the SLO framework rather than inherited it • Ability to walk through personally written SLIs and explain how they were negotiated with resistant teams • Strong Python • Comfortable reading and modifying Go or Rust • Practical experience with time-series and event telemetry at scale • Prometheus/OpenMetrics, Grafana, and an Alertmanager-class routing layer • Experience with ClickHouse or an equivalent columnar store for high-cardinality fleet data • Distributed systems debugging on bare metal and long-lived hosts • Configuration management and CI at production scale, including Ansible, GitLab CI, Jenkins, or close equivalents • Ability to design measurement for machines that cannot be owned or scraped, including push telemetry, sampling, clock skew, partial reporting, and customer-server privacy constraints • Strong written communication suitable for async work
• Professional development opportunities • Interesting and challenging projects • Mentor and knowledge-exchange programs • Fully remote work with flexible working hours • Work from any location worldwide • 24 days of paid vacation per year • 10 days of national holidays • Unlimited sick leave • Compensation for private medical insurance • Co-working reimbursement • Gym/sports reimbursement • Opportunity to receive a reward for the most innovative idea that the company can patent
Apply Now🔥 7 hours ago
Senior DevOps Engineer scaling AWS cloud infrastructure and CI/CD for an enterprise tech review platform trusted by Fortune 100 companies. Optimizing Ruby on Rails reliability and emerging AI/ML workloads.
AWS
Cloud
Docker
Google Cloud Platform
Kubernetes
Ruby
Ruby on Rails
Terraform
🔥 7 hours ago
Senior DevOps Engineer owning CI/CD and secure cloud infrastructure. Scaling content solutions that preserve theme-park and attraction memories.
AWS
Cloud
Distributed Systems
Docker
DynamoDB
EC2
Google Cloud Platform
JavaScript
Kubernetes
Linux
MongoDB
MySQL
Node.js
NoSQL
Python
SDLC
SQL
🔥 7 hours ago
Senior DevOps Engineer building secure, scalable cloud infrastructure for CI/CD pipelines and production services. Automating deployments, Kubernetes platforms, and Infrastructure-as-Code.
AWS
Cloud
Docker
Google Cloud Platform
Grafana
Kubernetes
Linux
Prometheus
Python
Shell Scripting
Terraform
Go
🕒 2 days ago
DevOps Engineer designing CI/CD automation and scalable Kubernetes environments for a growing DevOps team. Improving infrastructure security, reliability, and production operations in Poland.
Ansible
Azure
Docker
ElasticSearch
Grafana
Kubernetes
Linux
Python
Terraform
Vault
🕒 3 days ago
DevOps Engineer strengthening Inetum Polska’s Azure Landing Zone, governance, security, and networking. Supporting hybrid cloud operations and enterprise Azure adoption with Microsoft partners.
🇵🇱 Poland – Remote
💰 Post-IPO Equity on 2007-03
⏰ Full Time
🟡 Mid-level
🟠 Senior
⛑ DevOps & Site Reliability Engineer (SRE)
🗣️🇵🇱 Polish Required
Azure
Cloud
Kubernetes
Terraform