Senior Site Reliability Engineer

🔥 14 minutes ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Fortress Information Security

Fortress Information Security

201 - 500 employees

Founded 2015

🔒 Cybersecurity

🏛️ Government

🤖 Artificial Intelligence

💰 $125M Series C - Fortress Information Security on 2022-04

Cybersecurity • Government • Artificial Intelligence

Fortress Information Security is an AI-powered cybersecurity company that defends critical infrastructure, government agencies, and their supply chains against cyber threats and mission risk. The firm focuses on threat intelligence, vulnerability management, vendor and third‑party risk management (TPRM), supply chain and product cyber supply chain security (C-SCRM), SBOMs, and protecting public sector and critical infrastructure clients.

📋 Description

• Transition applications from traditional (non-containerized) Ansible deployments to containerized, orchestrated deployments in AWS and on-premises environments • Build upon current CI/CD efforts to support different deployment strategies (Blue/Green, Canary, etc.) • Support Development/QA/UAT efforts by building an environment for anyone in the company to test drive a release anywhere in the development lifecycle • Improve common infrastructure for developers, such as CI/CD pipelines, log/application monitoring, cluster management, and configuration management • Migrate application secrets and configuration from Ansible Vault to Hashicorp Vault • Automate infrastructure provisioning during deployment • Handle code deployments in all environments (cloud, on-premises) • Implement tools to monitor and alert with respect to service level metrics and objectives • Provide technical guidance and educate team members and coworkers on development and operations • Monitor relevant systems for availability and performance • Available to support daytime and after business hours release activities

🎯 Requirements

• 5-8 years hands-on experience in a SRE/DevOps role supporting production systems • Production experience supporting Linux-based infrastructure and administering services on AWS (RDS, VPC, ECR, CloudWatch, Cloud Formation, Lambda, API Gateway) and on-premises • Demonstrable experience with deployment technologies such as Kubernetes, Ansible, Jenkins and Terraform (or similar technologies) • Excellent documentation skills so anyone on the team can come up-to-speed on changes quickly • Excellent written and verbal communication skills • Experience implementing rolling upgrades (canary, blue/green, etc.) • Strong scripting and tooling skillset (Bash, Python, JS, etc.) • Self-motivated, resourceful and a persistent problem-solving aptitude with advanced time management skills • Ability to independently use and refine prompts to enhance the quality, efficiency, and insight of regular work processes • Must be willing to participate in technical interviews and technical questions, which may be recorded or transcribed for evaluation purposes (required)

🏖️ Benefits

• Remote and Hybrid working environment • Competitive pay structure • Medical, dental, vision plans with employees covered up to 90% with highly progressive options for dependents and families • Company paid life, short- and long-term disability insurance • Employee Assistance Program • 401(k) match • Flexible Paid Time Off • Parental Leave

Apply Now

Similar Jobs

🔥 12 hours ago

MyFitnessPal

51 - 200

🏥 Healthcare

🍽️ Food & Beverage

💼 Consulting

Site Reliability Engineer improving MyFitnessPal’s production systems reliability and security. Engaging in incident response, observability, and infrastructure management.

🕒 Yesterday

CXM

201 - 500

💸 Finance

💳 Fintech

Application Site Reliability Engineer focusing on .NET/C# services reliability for trading systems. Collaborating with software engineers to enhance service resilience and operational excellence.

🕒 Yesterday

NVIDIA

10,000+ employees

🏥 Healthcare

🏭 Manufacturing

🤖 Artificial Intelligence

DevOps Engineer supporting NVIDIA’s Rapids project for AI and data science initiatives. Collaborating with teams to ensure high-quality software releases and infrastructure maintenance.

🕒 Yesterday

Softgic

51 - 200

💼 Consulting

🔒 Cybersecurity

DevOps Specialist managing Google Cloud infrastructure at Softgic. Automating and optimizing cloud systems and collaborating with development teams.

🗣️🇪🇸 Spanish Required

🕒 Yesterday

Global Enterprise Services, LLC (GES)

11 - 50

💼 Consulting

📦 Logistics

Reliability Engineer responsible for cloud platform performance and incident response, managing compliance. Requires strong technical expertise and 8 years of experience.