Site Reliability Engineer

Job not on LinkedIn

🕒 August 31

🇺🇸 United States – Remote

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

👻 Ghost score 10%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Identiq

Identiq

51 - 200 employees

Founded 2018

💳 Fintech

🛍️ eCommerce

🔒 Cybersecurity

💰 $47M Series A on 2021-03

Fintech • eCommerce • Cybersecurity

Identiq is a company that specializes in payment optimization and customer identification solutions. They provide a private network designed to enhance payment acceptance rates, reduce fraud, and improve overall customer experiences without compromising sensitive data. Their technology enables risk-based decisions through the use of first-party data, ensuring that sensitive information remains secure and private throughout the validation process.

📋 Description

• Join the Engineering team as the first dedicated Site Reliability Engineer • Define what reliable means for production systems and establish the SRE practice from zero • Define SLIs and SLOs for core services • Implement Grafana dashboards and burn-rate-based alerting • Establish, own, and continuously improve incident management, including PagerDuty, on-call training, and incident command • Own the observability stack end to end across metrics, logs, traces, RUM, and synthetic checks • Partner with engineering teams to refine SLIs, SLOs, and error budgets and coach teams on SRE and observability best practices • Automate manual and repetitive operational work using infrastructure as code and tooling • Design and run load/performance tests and chaos engineering game days • Drive reliability and infrastructure projects independently at startup speed • Establish documented incident management from detection through blameless postmortems • Reduce alert noise and MTTR and improve confidence in reliability signals for release and investment decisions

🎯 Requirements

• Bachelor's degree in Computer Science, Computer Engineering, or equivalent formal training, with depth in operating systems, databases, and networking; a rigorous equivalent is accepted • Fundamental systems understanding required for diagnosing novel failures • Daily active use of AI tools to write and debug code, build dashboards and alerts, and increase execution speed • Demonstrated history of independently driving large, ambiguous reliability or infrastructure projects to completion, typically reflecting 5+ years in an SRE, DevOps, or production engineering role • Hands-on implementation of SLI/SLO/error-budget methodology • Strong experience with Grafana and PromQL • Experience with Grafana Alloy for Loki logs and a metrics backend such as Prometheus or Datadog • Experience with OpenTelemetry and a tracing/APM backend such as SigNoz, Uptrace, Tempo, Datadog, or New Relic • Experience with Real User Monitoring (RUM) and synthetic monitoring, such as Grafana Faro, Grafana Synthetic Monitoring, or k6 • Experience designing on-call rotations and incident command practices using PagerDuty or equivalent • Hands-on experience with load/performance frameworks such as Locust, k6, or JMeter • Experience with chaos engineering exercises • Proficiency in Python, Go, or Bash • Hands-on experience with Infrastructure as Code such as Terraform or Ansible • Hands-on experience with Kubernetes • Experience with at least one major cloud platform: AWS, GCP, or Azure • Ability to communicate technical root causes, tradeoffs, implementation details, mitigations, fixes, reliability status, risks, and priorities to technical and business stakeholders • Ability to work independently, resolve ambiguous problems quickly, take ownership, and manage multiple threads under time pressure

🏖️ Benefits

• Whole-person growth and personal and professional development • Energetic and collaborative environment • Excellent work/life balance • Medical insurance • Dental insurance • Vision insurance • Life insurance • 401k match • Paid time off (PTO) • Two office locations: Downtown Atlanta and Halcyon in Alpharetta

Apply Now

Similar Jobs

🕒 August 31

Akamai Technologies

5001 - 10000

🔒 Cybersecurity

Site Reliability Engineer improving reliability, performance, and scalability across Akamai’s distributed cloud and edge platform. Automating operations, strengthening observability, and leading incident response.

🇺🇸 United States – Remote

💵 $75.7k - $136.3k / year

💰 Post-IPO Equity on 2001-07

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🦅 H1B Visa Sponsor

infoinfo

🕒 August 31

Jabil

10,000+ employees

🚘 Automotive

🎖️ Defense

🏥 Healthcare

Lead cybersecurity, Fortinet, Arista, and site reliability engineering for Jabil’s manufacturing and test infrastructure. Secure networks, virtualization, storage, and production environments.

🇺🇸 United States – Remote

💵 $100.1k - $180.2k / year

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🦅 H1B Visa Sponsor

infoinfo

Cloud

Cyber Security

Switching

VMware

🕒 August 31

Peraton

10,000+ employees

💼 Consulting

🏥 Healthcare

📦 Logistics

Site Reliability Engineer maintaining AWS, GovCloud, and OpenShift production systems for Peraton’s national security missions. Automating deployments, observability, incident response, and infrastructure resilience.

🇺🇸 United States – Remote

💵 $104k - $166k / year

⏰ Full Time

🟠 Senior

🔴 Lead

⛑ DevOps & Site Reliability Engineer (SRE)

🦅 H1B Visa Sponsor

infoinfo

🕒 August 31

Peraton

10,000+ employees

💼 Consulting

🏥 Healthcare

📦 Logistics

Site Reliability Engineer maintaining reliable AWS and OpenShift production systems for Peraton, a national security and enterprise IT provider. Automating infrastructure, observability, deployments, and incident response.

🇺🇸 United States – Remote

💵 $104k - $166k / year

⏰ Full Time

🟠 Senior

🔴 Lead

⛑ DevOps & Site Reliability Engineer (SRE)

🦅 H1B Visa Sponsor

infoinfo

🕒 August 31

Rackner

11 - 50

💼 Consulting

🎖️ Defense

🤖 Artificial Intelligence

R&D DevSecOps Engineer building secure AI-enabled mission software and DevSecOps pipelines. Developing backend services, coding agents, and compliant delivery workflows for U.S. government defense customers.

🇺🇸 United States – Remote

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)