Search Remote Jobs

Staff Site Reliability Engineer

Job not on LinkedIn

🔥 3 minutes ago

🇺🇸 United States – Remote

💵 $177k - $240k / year

⏰ Full Time

🔴 Lead

⛑ DevOps & Site Reliability Engineer (SRE)

👻 Ghost score 3%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Fingerprint

Fingerprint

51 - 200 employees

Founded 2019

🔒 Cybersecurity

🔌 API

☁️ SaaS

💰 $32M Series B on 2021-11

Cybersecurity • API • SaaS

Fingerprint is a technology company focused on bot detection and identification solutions. The platform helps businesses ensure a secure online environment by accurately identifying users and distinguishing between real users and automated bots. They provide developer guides and workspace settings for easy integration into various applications, enhancing security and user experience.

📋 Description

• Serve as Fingerprint's first dedicated Site Reliability Engineer • Work with the Architect on platform design and with Cloud Platform on infrastructure • Partner with every product team on operating their systems • Define SLIs and SLOs for critical request paths and make them visible and actionable • Introduce and coach teams on error budgets • Own reliability metrics used by leadership • Strengthen incident detection, response, communication, postmortems, and follow-up • Improve alert quality, anomaly detection, escalation design, and shared tooling • Lead reliability reviews for high-risk changes and new services • Introduce game days and chaos exercises • Embed with teams on time-boxed reliability engagements • Develop Staff and Lead engineers as reliability leaders • Codify production-readiness, on-call, runbook, and change-safety practices • Partner with the Architect and tech leads to design reliability into systems • Investigate production incidents and write tooling, dashboards, and reference implementations • Lead AI adoption for incident investigation, postmortems, runbooks, observability, and safe AI-assisted operations • Report directly to the VP of Engineering

🎯 Requirements

• 10+ years of engineering experience • 3+ years as an SRE, production engineer, or reliability-focused Staff engineer operating across multiple teams • Experience owning reliability for a platform • Deep experience with SLI/SLO design and error budgets in practice • Experience driving product-team adoption of reliability practices • Experience leading incident response and postmortems for high-severity, customer-facing incidents • Hands-on knowledge of distributed-systems failure modes, including cache/database saturation, cascading failure, retry storms, capacity limits, degradation, and load shedding • Experience in high-throughput, low-latency environments • Fluency in Kubernetes, AWS, and modern observability tooling such as Datadog or equivalent • Ability to read and write production code in Go, TypeScript, or similar • Experience with infrastructure as code • Track record of leading through influence across teams • Experience coaching engineers to own reliability • Exceptional written communication and documented decision-making • Regular use of AI tools for incident investigation, telemetry analysis, runbooks, postmortems, and tooling • Ability to structure operational data for safe use by humans and AI agents • Pragmatic approach to balancing reliability, delivery, and risk • Must be authorized to work from the home location • Nice to have: experience in fraud detection, identity, payments, or other adversarial real-time domains • Nice to have: multi-region, cell-based, or failure-isolation architecture experience • Nice to have: Elasticsearch, Redis, DynamoDB, or Kafka at scale • Nice to have: familiarity with FinOps and cloud infrastructure reliability/cost trade-offs

🏖️ Benefits

• Pay transparency • Fully remote work arrangement • Ability to work from almost any country, subject to country restrictions • Inclusive work environment • No visa sponsorship required; teammates may work from their authorized home location

Apply Now

Similar Jobs

🕒 2 days ago

RealTime eClinical Solutions

51 - 200

🏥 Healthcare

🧬 Biotechnology

Principal DevOps Architect owning AWS, Terraform, CI/CD, observability, and AI/ML platforms. Ensuring HIPAA and SOC 2 compliance for clinical research SaaS.

🇺🇸 United States – Remote

💵 $155k - $195k / year

💰 Private Equity Round on 2022-01

⏰ Full Time

🔴 Lead

⛑ DevOps & Site Reliability Engineer (SRE)

🕒 3 days ago

Centex Technologies

51 - 200

💼 Consulting

📦 Logistics

📣 Marketing

DevOps Engineer 4 building secure CI/CD and cloud platforms for Centex Technologies’ programs and customers. Automating infrastructure, observability, security, and reliable software delivery.

🕒 3 days ago

Conga

1001 - 5000

☁️ SaaS

💸 Finance

🏢 Enterprise

Staff DevOps Engineer building secure cloud infrastructure and CI/CD systems for Conga’s commercial operations software. Managing Kubernetes, observability, IAM, cost optimization, and Terraform automation.

🕒 4 days ago

Avanade

10,000+ employees

💼 Consulting

📦 Logistics

📣 Marketing

Avanade manager architecting Azure DevOps, GitHub, and AI-enabled software delivery solutions for enterprise clients. Leading DevOps transformation, Copilot adoption, governance, and cloud engineering modernization.

🕒 4 days ago

Red Cell Partners

11 - 50

🏥 Healthcare

🎖️ Defense

💼 Consulting

Staff DevOps Engineer building secure CI/CD and deployment infrastructure for Red Cell’s federal technology companies. Supporting classified environments, ATO activities, and customer deployments.