Search Remote Jobs

NOC Engineer / SRE

🔥 2 minutes ago

🇬🇧 United Kingdom – Remote

⏰ Full Time

🟡 Mid-level

🟠 Senior

🌐 Network Operations

🇬🇧 UK Skilled Worker Visa Sponsor

infoinfo

👻 Ghost score 10%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of NICE

NICE

5001 - 10000 employees

Founded 1991

☁️ SaaS

🤖 Artificial Intelligence

📡 Telecommunications

SaaS • Artificial Intelligence • Telecommunications

NICE is a leading provider of AI-powered customer service automation solutions, transforming contact centers into world-class customer experience centers. Their CXone Mpower platform offers end-to-end automation of customer service workflows, integrating human and AI agents to deliver efficient and personalized customer interactions. NICE's offerings include AI for customer experience, digital and self-service solutions, workforce engagement and management, and complete cloud-based contact center platforms. They are recognized as a leader in the Contact Center as a Service (CCaaS) industry, providing tools for increased operational efficiency, employee engagement, and enhanced customer satisfaction.

📋 Description

• Act as a primary or escalation responder in a 24x7 on-call rotation • Lead or support Major Incident response, including triage, mitigation, and resolution • Coordinate across Engineering, Infrastructure, Security, and Product teams • Execute and improve runbooks, playbooks, and escalation paths • Drive blameless post-incident reviews and track corrective actions • Own service health monitoring across infrastructure, applications, and dependencies • Design and maintain alerting strategies aligned with SLIs/SLOs • Reduce alert fatigue through signal-to-noise improvements • Build dashboards using Grafana, Prometheus, Datadog, Splunk, and/or CloudWatch • Automate repetitive operational tasks and reduce manual toil • Improve mean time to detect and mean time to resolve • Develop scripts and tools in Python, Bash, Go, or similar • Implement self-healing and auto-remediation where possible • Partner with engineering teams to improve system reliability • Support and troubleshoot Linux systems, cloud platforms, and Kubernetes/containerized environments • Assist with capacity planning and availability reviews • Ensure operational readiness for production releases

🎯 Requirements

• Strong Linux systems administration • Experience with incident management and production support • Familiarity with cloud infrastructure, preferably AWS • Familiarity with Docker and Kubernetes • Familiarity with monitoring and alerting platforms • Scripting or programming experience in Python, Bash, Go, or similar • Understanding of networking fundamentals, including DNS, TCP/IP, and load balancing • Experience working in 24x7 NOC or production operations environments • Ability to handle high-pressure incidents calmly and effectively • Strong written and verbal communication for incident coordination • Comfort working from runbooks and improving them when needed • Experience defining or operating to SLOs/SLIs preferred • Prior migration from traditional NOC to SRE model preferred • Infrastructure as Code experience with Terraform, Ansible, or similar preferred • Exposure to security, compliance, or regulated environments preferred

Apply Now

Similar Jobs

🕒 June 18

TM Forum

51 - 200

💼 Consulting

📦 Logistics

📣 Marketing

Deputy Mission Lead overseeing Autonomous Network Operations at TM Forum, focusing on collaboration, strategy, and community building.

🇬🇧 United Kingdom – Remote

⏰ Full Time

🟠 Senior

🌐 Network Operations