Senior Site Reliability Engineer

🔥 0 minutes ago

🇺🇸 United States – Remote

💵 $150k - $187k / year

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

👻 Ghost score 3%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Blackpoint Cyber

Blackpoint Cyber

51 - 200 employees

💼 Consulting

🎖️ Defense

🔒 Cybersecurity

💰 $190M Series C on 2023-06

Consulting • Defense • Cybersecurity

Blackpoint Cyber is a technology-focused cybersecurity company headquartered in Maryland, USA. Established by former US Department of Defense and Intelligence security experts, Blackpoint leverages its real-world cyber experience to help Managed Service Providers (MSPs) safeguard their infrastructure and operations. The company offers a proprietary cybersecurity ecosystem, including its SNAP-Defense platform for Managed Detection and Response (MDR) services. Blackpoint's dedicated security analysts work 24/7 to combine various security measures, including network visualization and endpoint security, to monitor and respond to threats. Additionally, Blackpoint is launching LogIC, a logging and integrated compliance service designed to assist MSPs with cyber compliance requirements. The company's mission is to deliver comprehensive detection and response services to help MSPs combat the evolving threat landscape.

📋 Description

• Design, develop, and maintain highly scalable infrastructure using Infrastructure as Code (Terraform and Terragrunt) for automated cloud resource provisioning and orchestration • Own and optimize the AWS cloud environment for cost efficiency, security, and high availability • Manage and optimize Kubernetes clusters using Helm, ArgoCD, Istio, and Kustomize • Administer and scale Confluent Cloud and Apache Kafka data streaming infrastructure • Deploy, configure, and maintain Redis for caching and real-time data processing • Implement and maintain monitoring, alerting, and incident response frameworks using Prometheus, Grafana, Alert Manager, Grafana Cloud, OpsGenie/PagerDuty • Facilitate controlled feature deployments and progressive rollouts through LaunchDarkly/PostHog • Partner with software development teams to integrate new services, applications, and features into existing infrastructure • Diagnose and resolve complex system-level issues while maintaining performance and maximizing uptime • Drive continuous improvement of automation tooling, operational processes, and engineering methodologies • Stay current on emerging SRE trends and help the team adopt relevant industry advancements and best practices

🎯 Requirements

• 5+ years of experience in a Senior Site Reliability Engineer role or equivalent, with substantial emphasis on cloud infrastructure management and automation • Expertise in Infrastructure as Code (Terraform, Terragrunt) for enterprise-scale deployments • Comprehensive knowledge of AWS, including designing, implementing, and maintaining secure, scalable, resilient cloud architectures • Extensive hands-on experience with distributed data streaming (Confluent Cloud, Apache Kafka) • Proven experience with Redis for caching and Amazon RDS for relational database management • Experience with enterprise search and analytics platforms (OpenSearch, Elasticsearch, ChaosSearch) • Proficiency designing and implementing monitoring/alerting infrastructure (Prometheus, Grafana, Alert Manager, Grafana Cloud, OpsGenie/PagerDuty) • Practical experience with feature flag systems (LaunchDarkly/PostHog) for controlled release management • Extensive experience administering production-grade Kubernetes (Helm, ArgoCD, Istio); working knowledge of Kustomize • Strong problem-solving skills, with the ability to troubleshoot complex systems in production • Strong communication and collaboration skills, with experience working in Agile environments • Must be authorized to work in the United States without restriction • Must not require sponsorship to work within the United States • Nice to have: Experience with Terragrunt, Kafka, Google Cloud Platform, Microsoft Azure, cloud security/compliance, serverless computing, Jenkins, GitHub Actions, Node.js, Python, and/or Go

🏖️ Benefits

• Health Insurance • Vision Insurance • Dental Insurance • Life Insurance • Robust 401k plan • Discretionary Time Off • Other minor perks • Equity participation available to employees globally

Apply Now

Similar Jobs

🔥 47 minutes ago

First Due

201 - 500

☁️ SaaS

🏛️ Government

🤝 B2B

Platform SRE maintaining secure cloud infrastructure, CI/CD, and observability for First Due’s fire and EMS software. Improving reliability, deployments, scalability, and incident response.

🇺🇸 United States – Remote

💵 $165k / year

💰 $355M Private Equity Round - First Due on 2025-08

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🔥 1 hour ago

Fusable

201 - 500

🏗️ Construction

📦 Logistics

🛡️ Insurance

Sr. DevOps Engineer automating AWS infrastructure, containers, and CI/CD pipelines for Fusable’s trucking, agriculture, and construction data solutions. Managing secure, scalable cloud environments and production reliability.

🔥 3 hours ago

Environmental Management Authority

51 - 200

🏥 Healthcare

🛡️ Insurance

💼 Consulting

DevOps Engineer building Kubernetes, cloud, and deployment infrastructure for Ema’s agentic AI platform. Ensuring secure, observable, reliable scaling for enterprise workflows.

🕒 2 days ago

ICF

5001 - 10000

💼 Consulting

🏛️ Government

🏥 Healthcare

Lead DevOps Engineer operating multi-account AWS and Kubernetes platforms for ICF, a global advisory and technology services provider. Automating secure, observable enterprise application delivery.

🕒 2 days ago

Mastercam

201 - 500

🚘 Automotive

💼 Consulting

📦 Logistics

DevSecOps Security Engineer securing Mastercam’s cloud-native manufacturing software, infrastructure, and AI platforms. Integrating Kubernetes, CI/CD, identity, and application security controls.