Search Remote Jobs

Site Reliability Engineer

Job not on LinkedIn

šŸ”„ 0 minutes ago

šŸ‡ŗšŸ‡ø United States – Remote

ā° Full Time

🟔 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

šŸ‘» Ghost score 10%

infoinfo
Apply Now
Find Similar Remote Jobs

šŸ“Š Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Identiq

Identiq

51 - 200 employees

Founded 2018

šŸ’³ Fintech

šŸ›ļø eCommerce

šŸ”’ Cybersecurity

šŸ’° $47M Series A on 2021-03

Fintech • eCommerce • Cybersecurity

Identiq is a company that specializes in payment optimization and customer identification solutions. They provide a private network designed to enhance payment acceptance rates, reduce fraud, and improve overall customer experiences without compromising sensitive data. Their technology enables risk-based decisions through the use of first-party data, ensuring that sensitive information remains secure and private throughout the validation process.

šŸ“‹ Description

• Join the Engineering team as the first dedicated Site Reliability Engineer • Define what reliable means for production systems and establish the SRE practice from zero • Define SLIs and SLOs for core services • Implement Grafana dashboards and burn-rate-based alerting • Establish, own, and continuously improve incident management, including PagerDuty, on-call training, and incident command • Own the observability stack end to end across metrics, logs, traces, RUM, and synthetic checks • Partner with engineering teams to refine SLIs, SLOs, and error budgets and coach teams on SRE and observability best practices • Automate manual and repetitive operational work using infrastructure as code and tooling • Design and run load/performance tests and chaos engineering game days • Drive reliability and infrastructure projects independently at startup speed • Establish documented incident management from detection through blameless postmortems • Reduce alert noise and MTTR and improve confidence in reliability signals for release and investment decisions

šŸŽÆ Requirements

• Bachelor's degree in Computer Science, Computer Engineering, or equivalent formal training, with depth in operating systems, databases, and networking; a rigorous equivalent is accepted • Fundamental systems understanding required for diagnosing novel failures • Daily active use of AI tools to write and debug code, build dashboards and alerts, and increase execution speed • Demonstrated history of independently driving large, ambiguous reliability or infrastructure projects to completion, typically reflecting 5+ years in an SRE, DevOps, or production engineering role • Hands-on implementation of SLI/SLO/error-budget methodology • Strong experience with Grafana and PromQL • Experience with Grafana Alloy for Loki logs and a metrics backend such as Prometheus or Datadog • Experience with OpenTelemetry and a tracing/APM backend such as SigNoz, Uptrace, Tempo, Datadog, or New Relic • Experience with Real User Monitoring (RUM) and synthetic monitoring, such as Grafana Faro, Grafana Synthetic Monitoring, or k6 • Experience designing on-call rotations and incident command practices using PagerDuty or equivalent • Hands-on experience with load/performance frameworks such as Locust, k6, or JMeter • Experience with chaos engineering exercises • Proficiency in Python, Go, or Bash • Hands-on experience with Infrastructure as Code such as Terraform or Ansible • Hands-on experience with Kubernetes • Experience with at least one major cloud platform: AWS, GCP, or Azure • Ability to communicate technical root causes, tradeoffs, implementation details, mitigations, fixes, reliability status, risks, and priorities to technical and business stakeholders • Ability to work independently, resolve ambiguous problems quickly, take ownership, and manage multiple threads under time pressure

šŸ–ļø Benefits

• Whole-person growth and personal and professional development • Energetic and collaborative environment • Excellent work/life balance • Medical insurance • Dental insurance • Vision insurance • Life insurance • 401k match • Paid time off (PTO) • Two office locations: Downtown Atlanta and Halcyon in Alpharetta

Apply Now

Similar Jobs

šŸ”„ 34 minutes ago

Akamai Technologies

5001 - 10000

šŸ”’ Cybersecurity

Site Reliability Engineer improving reliability, performance, and scalability across Akamai’s distributed cloud and edge platform. Automating operations, strengthening observability, and leading incident response.

šŸ”„ 42 minutes ago

Jabil

10,000+ employees

🚘 Automotive

šŸŽ–ļø Defense

šŸ„ Healthcare

Lead Fortinet network security and site reliability engineering for Jabil’s manufacturing infrastructure. Securing networks, virtualization, storage, and production test environments.

šŸ”„ 2 hours ago

Everbridge

1001 - 5000

šŸ„ Healthcare

šŸ“¦ Logistics

šŸ’¼ Consulting

Senior Site Reliability Engineer II building resilient cloud platforms for Everbridge’s critical event management technology. Improving reliability, automation, observability, and incident response.

šŸ”„ 2 hours ago

CADRE GOVERNMENT SOLUTIONS

11 - 50

šŸ›ļø Government

šŸ’¼ Consulting

šŸ”’ Cybersecurity

DevSecOps Engineer designing AWS and Salesforce environments for government agencies. Automating infrastructure, CI/CD pipelines, testing environments, and VA-compliant security controls.

šŸ”„ 3 hours ago

Peraton

10,000+ employees

šŸ’¼ Consulting

šŸ„ Healthcare

šŸ“¦ Logistics

Site Reliability Engineer maintaining AWS, GovCloud, and OpenShift production systems for Peraton’s national security missions. Automating deployments, observability, incident response, and infrastructure resilience.