Site Reliability Architect

Job not on LinkedIn

🔥 0 minutes ago

🇺🇸 United States – Remote

💵 $136k - $175k / year

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🦅 H1B Visa Sponsor

info
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Summit

Summit

51 - 200 employees

Founded 2008

💼 Consulting

🏨 Hospitality

📣 Marketing

Consulting • Hospitality • Marketing

Summit is a SaaS platform that helps businesses create, customize, shorten and track URLs and QR Codes, build mobile-friendly landing pages, and analyze engagement across campaigns. It offers branded links, UTM tracking, dynamic QR codes, 2D barcodes, digital business cards, integrations (including Shopify and Canva), and a developer-facing API to automate and scale link management for teams and enterprises. Summit targets business customers with free and enterprise plans, focusing on marketing, customer experience, and analytics use cases.

📋 Description

• Define the observability and reliability architecture strategy across Summit’s platforms and services • Implement site reliability concepts based around SLOs, SLIs, and SLAs across teams and platforms • Partner with engineering and operations leadership to align system design with resilience and scalability goals • Lead the design, implementation, and governance of observability frameworks and standards • Oversee automation of monitoring, alerting, and incident response processes • Serve as escalation point and lead for complex, cross-platform incidents • Guide resolution and post-incident analysis • Evaluate and introduce emerging tools, frameworks, and practices • Mentor and coach engineering teams on reliability, automation, and continuous improvement • Deliver scalable observability frameworks, improved incident detection and resolution, reliability practices, and strategic platform roadmaps

🎯 Requirements

• Proven track record designing and implementing observability and reliability platforms at scale • Ability to influence architecture and strategy across engineering, operations, and product teams • Experience mentoring engineers and shaping organizational practices • Strategic thinking combined with hands-on problem-solving for complex systems • Expertise with observability stacks such as ELK, Grafana, Graphite, InfluxDB, LogicMonitor, or Prometheus • Advanced automation experience using Ansible, Terraform, or equivalent tooling • Experience designing reliability frameworks in multi-cloud environments, including Azure, AWS, or hybrid • Knowledge of Python, Go, Ruby, or JavaScript • Familiarity with compliance-heavy industries where reliability, security, and auditability are important

🏖️ Benefits

• Flexible Time Off • Medical, Dental & Vision coverage with HSA/HRA options • 401(k) with 4% match • Paid Parental Leave • Life & Disability Insurance • Wellness Support for mental and physical health • Free Colocation & Cloud Access • Work From Anywhere • Low-ego, get-it-done culture

Apply Now

Similar Jobs

🔥 48 minutes ago

MeridianLink

501 - 1000

💳 Fintech

🏦 Banking

☁️ SaaS

Senior Site Reliability Engineer operating MeridianLink’s serverless AWS platform. Managing production reliability, databases, backups, monitoring, incident response, and infrastructure automation.

🔥 2 hours ago

NVIDIA

10,000+ employees

🏥 Healthcare

🏭 Manufacturing

🤖 Artificial Intelligence

Senior SRE improving NVIDIA GeForce NOW’s reliable GPU cloud gaming infrastructure. Building observability, automation, Kubernetes, and incident-response tooling for service SLOs.

🔥 2 hours ago

SimpliGov

11 - 50

🏛️ Government

☁️ SaaS

⚡ Productivity

Senior DevOps/MLOps Engineer operating SimpliGov’s Azure AI platform for government customers. Building secure Kubernetes infrastructure, compliant inference paths, observability, releases, and cost controls.

🔥 3 hours ago

VetsEZ

201 - 500

🏥 Healthcare

💼 Consulting

📦 Logistics

Senior Backend DevOps Engineer operating AWS containerized microservices for the VA’s JLV clinical data viewer. Building CI/CD, observability, security, and disaster recovery capabilities.

🔥 3 hours ago

Group 1001

501 - 1000

💼 Consulting

🏥 Healthcare

💸 Finance

Senior Network Reliability Engineer automating insurance company network platforms at Group 1001. Applying SRE, cloud, Kubernetes, security, and observability practices to improve reliability and reduce operational toil.