IT Infrastructure Support Site Reliability Engineer II

🔥 21 minutes ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Astreya

Astreya

1001 - 5000 employees

Founded 2001

💼 Consulting

📦 Logistics

📣 Marketing

Consulting • Logistics • Marketing

Astreya is a leading global provider of IT Managed Services and Technology Solutions, known for its innovative approach to digital engineering and IT logistics. The company focuses on empowering businesses to excel in today's dynamic digital landscape by maximizing productivity and fostering innovation. Astreya offers a range of services including Data Center & Network Management, Digital Workplace Services, Next-Gen Digital Engineering, and Cybersecurity Services. With a commitment to excellence and a focus on operational frameworks, Astreya aims to transform technology into a valuable strategic asset for organizations worldwide.

📋 Description

• Partner with leadership to establish, monitor, and enforce Service Level Indicators (SLIs) and Service Level Objectives (SLOs) for infrastructure tooling, including configuration compliance rates, patch success rates, and deployment latency metrics. • Provide Level 3 expertise for tooling-specific incidents, focusing on automating incident remediation workflows and reducing Mean Time To Repair (MTTR) through intelligent automation and runbook development. • Identify and automate repetitive manual tasks across managed infrastructure, targeting measurable reductions in operational overhead (e.g., 50% reduction in manual server build time) through scripting and workflow automation. • Conduct thorough root cause analysis and lead blameless postmortems for all major service-impacting incidents, driving systemic improvements in tooling reliability and infrastructure resilience. • Engineer and maintain automated processes and scripts to populate, update, and synchronize asset management platforms (e.g., NetBox), configuration management databases, and monitoring systems for internal and external stakeholders. • Design, develop, and deploy full-stack applications, custom plugins, and automation scripts to extend functionality of management and monitoring systems, enabling direct device interaction for configuration management. • Develop and maintain fully automated Infrastructure-as-Code configurations for Windows and Linux server roles using tools such as Ansible, Terraform, or Puppet, including drift detection and auto-remediation capabilities. • Build end-to-end automation pipelines for vulnerability patching, security baseline enforcement (CIS benchmarks), and continuous compliance auditing against internal and regulatory standards for physical security devices. • Develop API-driven tools for network configuration management, automated firmware updates, zero-touch provisioning, pre/post-change validation, and real-time network health monitoring across the device fleet. • Deploy and standardize monitoring agents, centralized log collection systems, and custom dashboards with alerts based on critical SLIs (latency, error rate, saturation, traffic) for servers and edge devices. • Build and maintain custom monitoring exporters for the physical security device fleet, including camera systems, ensuring accurate metrics and structured, multi-severity logging output. • Build diagnostic tooling to correlate timestamps across distributed log streams and detect clock drift or NTP desync, a recurring root cause of false-positive outages across the device fleet. • Build automation scripts for intelligent ticket handling, problem validation, and escalation workflows within enterprise ticketing systems, ensuring 2-hour initial response SLAs are consistently met. • Support foundational security improvements across the device fleet, including managed credential/access controls and automated configuration backup. • Participate in 24x5 on-call rotation to provide timely support for infrastructure systems, security devices, and related tooling, ensuring service continuity and rapid incident response.

🎯 Requirements

• 6+ years of experience in Infra Automation Engineering, or Infrastructure Engineering. • Strong proficiency in Python, Bash, and PowerShell for automation scripting, with experience in Go for building high-performance backend services and APIs. • Hands-on experience with Infrastructure-as-Code tools (Terraform, Ansible, Chef, or Puppet) and configuration management practices, including drift detection, version control, and automated remediation. • Advanced knowledge of Linux and Windows server environments, including Tier 3 troubleshooting capabilities, system hardening, and enterprise-scale server management. • Solid understanding of enterprise networking concepts, Cisco device administration, network automation protocols (NETCONF/RESTCONF), and experience with network monitoring and flow analysis tools. • Experience implementing and managing monitoring solutions (Prometheus, Grafana, Datadog) or comparable proprietary internal monitoring and metrics-streaming systems (e.g., Monarch, Streamz), and centralized logging platforms (ELK Stack), with ability to create custom dashboards and alerting rules. • Experience deploying and customizing a CMDB/IPAM platform (e.g., NetBox) as a source of truth for device inventory and downstream automation. • Comfort operating within a large-scale, cloud-hosted enterprise environment (Kubernetes, Terraform, Helm), including familiarity with internal development and code-review tooling (e.g., Cider, Critic) or comparable large-scale internal toolchains. • Experience writing and maintaining custom monitoring exporters/agents for edge/IoT and physical security devices, including structured, glog-style logging output. • Proficiency in advanced text-processing and scripting (e.g., awk/gawk) for log parsing and timestamp correlation, with working knowledge of NTP/clock synchronization practices.

Apply Now

Similar Jobs

🕒 Yesterday

Sardine

51 - 200

🔒 Cybersecurity

📋 Compliance

💳 Fintech

DevOps Engineer at Sardine improving infrastructure and tooling for a remote-first financial crime platform. Collaborating to ensure reliable, scalable, and cost-efficient systems.

AWS

Cloud

Distributed Systems

Google Cloud Platform

Kubernetes

Prometheus

Python

Terraform

Go

🕒 Yesterday

Holafly

501 - 1000

✈️ Travel

📡 Telecommunications

👥 B2C

DevSecOps Engineer at Holafly, designing secure GCP foundations and automating with Terraform. Safeguarding connectivity for millions of travelers through infrastructure and security enhancements.

Ansible

Cloud

Docker

Google Cloud Platform

Kubernetes

Python

Terraform

🕒 4 days ago

Arista Networks

1001 - 5000

🏢 Enterprise

📡 Telecommunications

Site Reliability Engineer maintaining scalable systems and collaborating with development teams at Arista Networks. Implementing best practices and operational excellence in a hybrid cloud environment.

Linux

Python

Shell Scripting

Unix

Go

🕒 July 17

Arista Networks

1001 - 5000

🏢 Enterprise

📡 Telecommunications

Site Reliability Engineer at Arista Networks managing scalable cloud infrastructure and internal user base. Collaborating with engineering teams to enhance developer experience and workflow efficiency in a hybrid environment.

Linux

Python

Shell Scripting

Unix

Go

🕒 July 2

Arista Networks

1001 - 5000

🏢 Enterprise

📡 Telecommunications

Site Reliability Engineer at Arista Networks maintaining and supporting critical production systems. Collaborating with product development teams to enhance scaling, reliability, and workflow efficiency.

Linux

Python

Shell Scripting

Unix

Go