Senior Site Reliability Engineer – AWS, EKS

Job not on LinkedIn

🔥 0 minutes ago

☕ Washington – Remote

infoinfo

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

👻 Ghost score 25%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Salve.Inno

Salve.Inno

11 - 50 employees

Founded 2024

💼 Consulting

📣 Marketing

📦 Logistics

Consulting • Marketing • Logistics

Salve. Inno is a recruitment and consulting firm that connects exceptional talent with businesses through personalized hiring strategies and global remote sourcing. The company specializes in recruitment for roles across sectors such as marketing, forex, and iGaming, offering candidate sourcing, screening, and career-site driven hiring experiences while emphasizing DE&I, communication, and innovative process building. Founded in 2024 and headquartered in Gdańsk, Poland, Salve. Inno operates with a small team and a global footprint via remote job listings and consulting services.

📋 Description

• Own the reliability, availability and operational health of production services running on AWS and Amazon EKS • Operate and troubleshoot Kubernetes clusters in production, including lifecycle, upgrades, node management, networking, scaling, capacity and workload reliability • Participate actively in on-call and pager rotations and take ownership of production incidents • Lead or play a key technical role during P1/P2 and Sev1/Sev2 incidents, including diagnosis, mitigation, recovery and communication • Coordinate technical incident bridges and communicate directly with customers during production escalations • Perform root cause analysis and lead blameless postmortems • Define, monitor and improve SLIs, SLOs and error budgets • Develop and maintain actionable alerts, operational runbooks and automated remediation • Build and improve infrastructure using Terraform/Terragrunt and Infrastructure as Code practices • Operate GitOps-based delivery environments using Argo CD or FluxCD • Improve Kubernetes scaling and efficiency using Karpenter, KEDA and native Kubernetes autoscaling • Build and improve observability using Prometheus, Grafana, OpenTelemetry, Datadog and/or ELK • Support highly available distributed and event-driven systems, including Kafka/MSK environments • Design, implement and validate disaster recovery and business continuity mechanisms against RTO and RPO objectives • Identify recurring operational problems and eliminate toil through automation and engineering • Improve AWS performance, scalability, security and cost efficiency • Collaborate with software, platform and engineering teams to build reliability throughout the development lifecycle • Contribute to continuous improvement of incident management, operational readiness and SRE practices • Interact directly with customers during technical discussions and production escalations

🎯 Requirements

• Significant professional experience as a hands-on Site Reliability Engineer, Production Engineer or senior Platform Engineer with direct production ownership • Several years of recent, hands-on experience operating production environments on AWS • Strong, demonstrable experience operating Amazon EKS in production • Deep Kubernetes operational knowledge, including cluster administration, upgrades, nodes, autoscaling, networking, troubleshooting and production failure scenarios • Proven participation in a production on-call/pager rotation • Demonstrable ownership of significant production incidents, including troubleshooting, mitigation, recovery, RCA and post-incident improvements • Practical experience with SLIs, SLOs, error budgets, alerting and runbooks • Strong Infrastructure as Code experience with Terraform and/or Terragrunt • Production experience with Kubernetes delivery and GitOps practices; Argo CD or FluxCD strongly preferred • Strong production observability experience with Prometheus, Grafana, OpenTelemetry, Datadog or ELK • Experience operating highly available, distributed production systems • Strong understanding of AWS networking, IAM, security, availability and resilience • Experience implementing and testing disaster recovery strategies with measurable RTO/RPO objectives • Proven external customer-facing technical experience • Ability to explain complex technical problems clearly and make sound decisions during high-pressure production incidents • Strong troubleshooting mindset and ability to work independently during complex production failures • Strong professional English communication skills (minimum C1) for regular interaction with clients • Coherent track record demonstrating sustained hands-on production engineering ownership

🏖️ Benefits

• Full-time permanent B2B cooperation • Fully remote working environment • Senior hands-on engineering position with meaningful ownership of business-critical production systems • Opportunity to work on complex AWS, Kubernetes and distributed-system environments at scale • Direct influence over reliability engineering, operational practices and platform improvements • Modern engineering environment with strong emphasis on automation, observability and continuous improvement • Collaboration with experienced engineering, platform and product teams • Opportunity to introduce and use modern approaches, including AI-assisted engineering and operational automation • Long-term opportunity for engineers who want to remain deeply technical and close to production • Inclusive, respectful workplace regardless of gender, ethnicity, or background

Apply Now

Similar Jobs

🔥 10 minutes ago

11:11 SYSTEMS

201 - 500

💼 Consulting

🏥 Healthcare

📦 Logistics

Cloud Operations Engineer managing VMware infrastructure for 11:11’s cloud compute systems. Handling provisioning, capacity planning, monitoring, escalations, and customer support.

🔥 4 hours ago

Koniag Government Services

1001 - 5000

🏛️ Government

🎖️ Defense

💼 Consulting

Senior AWS DevOps Engineer building AWS, Kubernetes, CI/CD, and observability platforms. Supporting Koniag’s federal government customers with secure, reliable digital services.

🔥 6 hours ago

Weekday (YC W21)

11 - 50

💼 Consulting

👥 HR Tech

☁️ SaaS

Cloud/DevOps Engineer evaluating Kubernetes, AWS, IaC, and CI/CD solutions for GenAI infrastructure. Creating reference tasks and feedback to improve AI training and inference systems.

🔥 11 hours ago

Guild Mortgage

1001 - 5000

🏗️ Construction

💸 Finance

🏠 Real Estate

DevOps Manager leading engineers, CI/CD, cloud platforms, and infrastructure automation for Guild Mortgage’s mortgage banking operations. Driving secure, predictable technology delivery across enterprise initiatives.

🔥 13 hours ago

Leidos

10,000+ employees

🏥 Healthcare

💼 Consulting

📦 Logistics

DevSecOps Engineer securing Leidos defense systems and JADC2 infrastructure. Automating CI/CD, cyber remediation, and compliant systems administration for DoD mission platforms.