Senior Site Reliability Engineer

🔥 5 minutes ago

🇺🇸 United States – Remote

💵 $152k - $205k / year

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

👻 Ghost score 5%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Fingerprint

Fingerprint

51 - 200 employees

Founded 2019

🔒 Cybersecurity

🔌 API

☁️ SaaS

💰 $32M Series B on 2021-11

Cybersecurity • API • SaaS

Fingerprint is a technology company focused on bot detection and identification solutions. The platform helps businesses ensure a secure online environment by accurately identifying users and distinguishing between real users and automated bots. They provide developer guides and workspace settings for easy integration into various applications, enhancing security and user experience.

📋 Description

• Own the reliability of core production systems end to end • Define and maintain SLIs and SLOs, dashboards, alerts, and error-budget practices • Improve alert quality and anomaly/correctness detection • Lead incident response, restore service, and write actionable postmortems • Build secure, resilient, and cost-efficient infrastructure with explicit failure-mode handling • Perform load testing, profiling, saturation analysis, and capacity planning • Improve change safety through progressive delivery, automated rollback, pre-production signals, and safe deployment practices • Manage infrastructure through code and configuration, primarily using Terraform • Design, write, and ship software and developer-facing tooling • Run game days and chaos exercises • Partner with product engineering teams on production readiness, capacity, failure modes, rollback plans, runbooks, and on-call handoff • Participate in and improve the on-call rotation • Apply a security lens to engineering work and peer reviews • Serve as the go-to person for difficult production problems and mentor engineers through code review, pairing, and design feedback

🎯 Requirements

• 6–10 years of experience in SRE, production engineering, infrastructure, or backend engineering within primarily cloud-based environments (AWS preferred) • Track record of owning a system end to end • Hands-on experience defining and operating against SLIs, SLOs, and error budgets • Experience leading or serving as a primary responder on high-severity, customer-facing incidents • Depth in distributed-systems failure modes in high-throughput, low-latency environments • Depth in cloud infrastructure fundamentals, including networking, load balancing, containerization (EKS/Kubernetes), and distributed systems • Strong hands-on experience managing infrastructure through code and configuration (Terraform or equivalent) • Solid programming skills in Go, Python, or a comparable language • Fluency with observability tooling such as Datadog, Prometheus, Grafana, or OpenTelemetry • Hands-on experience operating Redis/ElastiCache in production, including cluster/shard management, failover behavior, memory eviction policies, and scaling strategies • Fluency with software engineering best practices, including source control, code review, comprehensive test coverage, and safe deployment • High level of personal ownership and autonomy, with experience working without clearly defined requirements • Pragmatism in balancing reliability and delivery • Strong written and verbal communication in English • AI-native use of AI tools for incident investigation, telemetry analysis, runbooks, and tooling • Must be authorized to work from the home location • Visa sponsorship is not provided

🏖️ Benefits

• 100% remote work • Ability to join the workforce from almost any country, subject to country restrictions • Visa sponsorship is not provided • Inclusive work environment • CCPA and GDPR notices for applicable residents

Apply Now

Similar Jobs

🔥 4 hours ago

Impiricus

11 - 50

🏥 Healthcare

💼 Consulting

🍽️ Food & Beverage

DevOps engineer building scalable AWS infrastructure and CI/CD for Impiricus’s AI-powered HCP engagement platform. Improving reliability, security, observability, and QA automation.

🇺🇸 United States – Remote

💵 $100k - $130k / year

💰 $3M Seed Round on 2022-04

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🔥 5 hours ago

DecisionPoint Corporation

51 - 200

🎖️ Defense

💼 Consulting

📦 Logistics

DevSecOps technician maintaining AWS cloud environments and deployment platforms for USTRANSCOM. Supporting secure transportation and logistics applications through Kubernetes, GitLab, and ArgoCD.

🔥 5 hours ago

Ondo Finance

51 - 200

₿ Crypto

💳 Fintech

💸 Finance

Site Reliability Engineer owning reliability, observability, and performance for Ondo Finance’s blockchain-enabled trading platform. Operating Go/Rust services and multi-region AWS Kubernetes infrastructure.

🔥 6 hours ago

Peraton

10,000+ employees

💼 Consulting

🏥 Healthcare

📦 Logistics

Site Reliability Engineer operating Peraton’s AWS, GovCloud, and ROSA production infrastructure. Improving observability, incident response, resilience, and deployment automation for national security systems.

🔥 6 hours ago

Peraton

10,000+ employees

💼 Consulting

🏥 Healthcare

📦 Logistics

Site Reliability Engineer operating Peraton’s AWS and OpenShift production infrastructure. Managing reliability, observability, incident response, releases, and infrastructure automation for national security missions.