IT Infrastructure Support, Site Reliability Engineer II

🔥 0 minutes ago

🇮🇪 Ireland – Remote

💵 €50.6k - €63.3k / year

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

👻 Ghost score 0%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Astreya

Astreya

1001 - 5000 employees

Founded 2001

💼 Consulting

📦 Logistics

📣 Marketing

Consulting • Logistics • Marketing

Astreya is a leading global provider of IT Managed Services and Technology Solutions, known for its innovative approach to digital engineering and IT logistics. The company focuses on empowering businesses to excel in today's dynamic digital landscape by maximizing productivity and fostering innovation. Astreya offers a range of services including Data Center & Network Management, Digital Workplace Services, Next-Gen Digital Engineering, and Cybersecurity Services. With a commitment to excellence and a focus on operational frameworks, Astreya aims to transform technology into a valuable strategic asset for organizations worldwide.

📋 Description

• Ensure the reliability, scalability, and performance of critical physical security infrastructure, including IP camera fleets, access control systems, Cisco switches, servers, networks, and cloud environments • Build and maintain automation tools, a centralized CMDB, monitoring systems, and enterprise-grade infrastructure management processes • Partner with leadership to establish, monitor, and enforce infrastructure SLIs and SLOs • Provide Level 3 expertise for tooling-specific incidents • Automate incident remediation workflows and develop runbooks to reduce MTTR • Identify and automate repetitive manual tasks through scripting and workflow automation • Conduct root cause analyses and lead blameless postmortems for major service-impacting incidents • Engineer automated processes and scripts for asset management platforms, CMDBs, and monitoring systems • Design, develop, and deploy full-stack applications, custom plugins, and automation scripts for management and monitoring systems • Maintain Infrastructure-as-Code configurations for Windows and Linux server roles, including drift detection and auto-remediation • Build automation pipelines for vulnerability patching, CIS security baseline enforcement, and continuous compliance auditing • Develop API-driven tools for network configuration management, firmware updates, zero-touch provisioning, validation, and network health monitoring • Deploy and standardize monitoring agents, centralized logging, dashboards, and alerts • Build custom monitoring exporters for physical security devices and camera systems • Develop diagnostic tooling for distributed log timestamp correlation and clock-drift/NTP-desynchronization detection • Build automation scripts for ticket handling, problem validation, and escalation workflows while meeting 2-hour initial response SLAs • Support managed credential/access controls and automated configuration backups • Participate in a 24x5 on-call rotation to ensure service continuity and rapid incident response • Collaborate with cross-functional teams to define service objectives, reduce operational toil, and improve system resilience

🎯 Requirements

• 6+ years of experience in Infra Automation Engineering or Infrastructure Engineering • Strong proficiency in Python, Bash, and PowerShell • Experience with Go for building high-performance backend services and APIs • Hands-on experience with Terraform, Ansible, Chef, or Puppet • Configuration management experience, including drift detection, version control, and automated remediation • Advanced knowledge of Linux and Windows server environments • Tier 3 troubleshooting capabilities • System hardening and enterprise-scale server management experience • Understanding of enterprise networking concepts and Cisco device administration • Experience with NETCONF/RESTCONF, network monitoring, and flow analysis tools • Experience with Prometheus, Grafana, Datadog, Monarch, Streamz, or comparable monitoring systems • Experience with centralized logging platforms such as ELK Stack • Ability to create custom dashboards and alerting rules • Experience deploying and customizing a CMDB/IPAM platform such as NetBox • Comfort operating in a large-scale, cloud-hosted enterprise environment • Familiarity with Kubernetes, Terraform, and Helm • Familiarity with internal development and code-review tooling such as Cider and Critic, or comparable toolchains • Experience writing and maintaining custom monitoring exporters/agents for edge/IoT and physical security devices • Structured, glog-style logging output experience • Proficiency in advanced text processing and scripting such as awk/gawk • Working knowledge of NTP/clock synchronization practices • Availability for 24x5 support and on-call rotation

🏖️ Benefits

• Base salary of €50,640–€63,300 gross annually • Performance-based bonuses may apply • Benefits-related payments may apply • Other general incentives may apply • Learning, collaboration, and career growth opportunities • Inclusive culture with diverse perspectives and opportunities to make an impact

Apply Now

Similar Jobs

🕒 August 7

Arista Networks

1001 - 5000

🏢 Enterprise

📡 Telecommunications

Site Reliability Engineer operating secure, scalable infrastructure for Arista Networks’ cloud networking products. Automating CI/CD, monitoring, incident response, and developer productivity systems.

Ansible

Docker

ElasticSearch

Grafana

Java

Linux

MariaDB

MongoDB

Postgres

Prometheus

Python

Shell Scripting

Spinnaker

Unix

Go

🕒 August 5

Twilio

5001 - 10000

🔌 API

🤝 B2B

DevOps Engineer shaping Twilio’s OpenTelemetry-first observability platform. Architecting scalable telemetry systems, developer tooling, and APIs for reliable, cost-effective operations.

AWS

Cloud

Distributed Systems

Grafana

Java

Kafka

Kubernetes

Prometheus

Python

Go

🕒 August 4

Red Hat

10,000+ employees

🏢 Enterprise

Senior Site Reliability Engineer modernizing Red Hat's internal enterprise open-source applications and infrastructure. Automating, securing, and improving highly available cloud and container platforms.

Ansible

Cloud

Docker

Kubernetes

Linux

OpenShift

Prometheus

Python

SDLC

Terraform

🕒 August 4

Arista Networks

1001 - 5000

🏢 Enterprise

📡 Telecommunications

Senior Site Reliability Engineer building and operating scalable, secure production infrastructure for Arista Networks’ cloud networking platforms. Automating operations, improving observability, and ensuring reliable deployments from Ireland.

AWS

Azure

Cloud

Distributed Systems

Docker

Google Cloud Platform

Grafana

Kubernetes

Linux

Postgres

Prometheus

Python

Shell Scripting

Spinnaker

Terraform

Unix

Go

🕒 August 3

Arista Networks

1001 - 5000

🏢 Enterprise

📡 Telecommunications

Site Reliability Engineer at Arista Networks focusing on building and operating critical production systems for scalability and reliability. Engaging in a collaborative remote role from Ireland.

AWS

Azure

Cloud

Distributed Systems

Docker

Google Cloud Platform

Grafana

Kubernetes

Linux

Postgres

Prometheus

Python

Shell Scripting

Spinnaker

Terraform

Unix

Go