Cloud Reliability Engineer – Recovery

Job not on LinkedIn

🕒 May 5

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of AlphaSense

AlphaSense

1001 - 5000 employees

Founded 2011

💼 Consulting

🏥 Healthcare

📣 Marketing

💰 Debt Financing on 2022-06

Consulting • Healthcare • Marketing

AlphaSense is a market intelligence and search platform that empowers companies to unlock critical insights across an extensive universe of public and private content, including company filings, broker research, expert calls, and market trends. Its AI-driven technology affords users the ability to conduct comprehensive due diligence with speed and accuracy, reducing uncertainty and blind spots in decision-making. Trusted by major corporations, financial institutions, and asset management firms globally, AlphaSense serves sectors such as financial services, health care, and technology by providing generative AI-powered solutions that integrate internal proprietary data with external premium content.

📋 Description

• Design and implement multi-region, multi-AZ AWS architectures that meet RTO/RPO targets • Engineer active-active and active-passive failover patterns using Route 53, Global Accelerator, and CloudFront • Build automated DR runbooks and playbooks using AWS Systems Manager Automation and Step Functions • Implement chaos engineering practices using AWS Fault Injection Simulator (FIS) to validate resiliency • Architect cross-region replication strategies for S3, DynamoDB Global Tables, RDS, and Aurora Global • Review containerized workloads using Kubernetes, ensuring resilience through self-healing, auto-scaling, and multi-cluster or multi-region deployments. • Administer AWS Backup across all services (EC2, EBS, RDS, EFS, FSx, DynamoDB, Aurora) with policy-based automation • Design immutable backup vaults and cross-account/cross-region backup replication pipelines • Develop and automate data recovery testing procedures, ensuring integrity and meeting defined SLAs • Implement point-in-time recovery (PITR) for databases and storage; validate via regular restore drills • Maintain Business Continuity Plans (BCP) and Disaster Recovery (DR) strategies, including tracking RTO (Recovery Time Objective) and RPO (Recovery Point Objective).

🎯 Requirements

• 5+ years in cloud infrastructure, SRE, or IT disaster recovery engineering roles • 3+ years of hands-on AWS experience in production environments at scale • Proven delivery of multi-region DR architectures with defined and tested RTO/RPO targets • Expert-level proficiency with core AWS resilience services (see skills matrix below) • Strong scripting skills: Python, Bash, or PowerShell for automation and orchestration • Experience with Infrastructure as Code: Terraform and/or AWS CloudFormation • Solid understanding of networking fundamentals: VPC, TGW, Direct Connect, VPN, DNS failover • Excellent written and verbal communication; able to produce executive-level DR reports.

🏖️ Benefits

• Health insurance • Retirement plans • Paid time off • Flexible work arrangements • Professional development opportunities

Apply Now

Similar Jobs

🕒 April 28

K&A Engineering

1 - 10

🏗️ Construction

⚡ Energy

📋 Compliance

Senior Transmission Line Engineer designing overhead transmission lines at K&A Engineering. Collaborating with teams, producing engineering deliverables, and providing technical oversight in India.

🕒 April 17

Sophos

1001 - 5000

💼 Consulting

🏥 Healthcare

🏭 Manufacturing

Detection Engineer responsible for analyzing advanced security threats. Collaborating with teams to translate threat intelligence into high-fidelity detections for Sophos.

Firewalls

Linux

Numpy

Pandas

Python

Unix

🕒 April 16

Exavalu

201 - 500

💼 Consulting

🏥 Healthcare

🛡️ Insurance

Microsoft Purview Engineer role requiring Azure expertise and 3+ years of experience in data governance. Join a dynamic team at Exavalu with a focus on digital transformation.

Azure

ETL

SQL

🕒 April 14

Smart Working

51 - 200

💼 Consulting

🏥 Healthcare

📣 Marketing

Fabric Engineer at Smart Working creating data solutions with Microsoft Fabric for business data needs. Collaborating with cross-functional teams to enhance data visibility and reliability.

Azure

Cloud

Python

SQL

🕒 April 9

ClickHouse

51 - 200

☁️ SaaS

🏢 Enterprise

🤖 Artificial Intelligence

Senior Consulting Engineer providing consultative support to ClickHouse customers. Collaborating with global teams for technical oversight and engineering assistance in a remote environment.

AWS

Azure

Cloud

Distributed Systems

EC2

Google Cloud Platform

Kubernetes

Linux

Unix