Search Remote Jobs

Senior Site Reliability Engineer

Job not on LinkedIn

🔥 0 minutes ago

🇺🇸 United States – Remote

đź’µ $130k - $170k / year

⏰ Full Time

đźź  Senior

⛑ DevOps & Site Reliability Engineer (SRE)

đź‘» Ghost score 0%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Climavision

Climavision

11 - 50 employees

đź’Ľ Consulting

📦 Logistics

🤖 Artificial Intelligence

đź’° $100M Series A on 2021-06

Consulting • Logistics • Artificial Intelligence

Climavision is a pioneering weather intelligence company that revolutionizes weather forecasting through its advanced suite of AI-powered products and services. With a focus on precision and customization, Climavision offers hyper-accurate weather insights for various industries, including agriculture, energy, and transportation. Their innovative Horizon AI models integrate diverse data sources, enabling businesses and government agencies to make informed decisions by anticipating weather-related risks and optimizing operations.

đź“‹ Description

• Own production reliability for customer-facing platform and data services across Azure, colocation, and edge Kubernetes environments. • Support the shared SRE function across the radar network and weather intelligence business. • Define and improve SLIs, SLOs, alerting standards, and operational metrics. • Build and own fleet observability through shared dashboards and proactive alerting. • Design and build automated recovery and self-healing for production systems. • Optimize cluster resources and costs, including right-sizing workloads and nodes and migrating workloads off Azure. • Coordinate production incident response, including troubleshooting, mitigation, communication, and postmortem analysis. • Diagnose and resolve complex issues across application services, Kubernetes infrastructure, storage, and distributed systems. • Drive multi-replica and multi-cluster high availability, including workload placement, scheduling, failover, traffic routing, and data replication. • Operate and improve self-managed Kubernetes platforms across cloud-hosted, colocation, and edge clusters. • Execute Kubernetes upgrades, patching, cluster health, node management, and production change management. • Improve observability, autoscaling, ingress, distributed storage, resiliency, and operational performance. • Design and validate Kubernetes workloads for resiliency, scalability, efficiency, and graceful degradation. • Partner with software engineering teams on production readiness, deployment safety, and operational visibility. • Maintain deployment pipelines, Helm charts, Kubernetes manifests, and infrastructure automation. • Support metrics, logging, distributed tracing, dashboarding, and alerting platforms. • Conduct performance engineering and capacity planning for peak weather-event demand. • Facilitate blameless postmortems and complete operational follow-up items. • Improve disaster recovery, failover, and business continuity capabilities. • Drive automation, toil reduction, game days, production readiness reviews, and reliability best practices. • Mentor teams on reliability engineering and production operations practices. • Participate in rotating weekday and weekend on-call schedules.

🎯 Requirements

• A bachelor's degree in computer science, software engineering, or a related field; equivalent professional experience considered. • Minimum of 7 years of experience in Site Reliability Engineering, DevOps, Production Engineering, Platform Engineering, or a related infrastructure-focused role, with at least 4 years in a role formally titled Site Reliability Engineer or carrying explicit SLO / error-budget accountability. • Deep, hands-on experience operating native Kubernetes; managed distributions such as AKS and EKS are acceptable, but native or self-managed Kubernetes is strongly preferred and is the primary technical requirement. • Experience optimizing Kubernetes clusters, including right-sizing workloads and node pools, resource management, and reducing infrastructure cost without sacrificing reliability. • Experience increasing operational visibility, including building dashboards, metrics pipelines, and alerting. • Experience designing and operating workloads for safe horizontal scaling across multiple replicas, including idempotency, concurrency, and state handling. • Experience designing or operating multi-cluster high-availability architectures, including failover behavior, traffic routing, and cross-cluster service deployment. • Experience supporting customer-facing production systems with uptime, reliability, and incident-response responsibilities. • Experience diagnosing and resolving production incidents across application, platform, and Kubernetes infrastructure layers. • Experience operating Kubernetes outside strictly managed cloud environments, including bare-metal, colocation, edge, or hybrid infrastructure. • Experience with Kubernetes operational tooling and ecosystem technologies such as Rancher, Helm, autoscaling frameworks, observability stacks, or distributed storage systems. • Strong understanding of infrastructure automation and Infrastructure as Code using tools such as Terraform and Ansible. • Experience supporting CI/CD and production deployment pipelines; GitHub Actions is used for CI/CD. • Experience with monitoring, logging, and observability platforms such as DataDog, Prometheus, Grafana, Loki, OpenTelemetry, or comparable technologies. • Experience operating distributed systems and microservice-based architectures in production. • Working knowledge of Microsoft Azure infrastructure. • Strong troubleshooting skills across infrastructure, application, and platform layers. • Experience participating in a structured production on-call rotation supporting business-critical systems. • Working familiarity with Jira, Confluence, and Microsoft Entra. • Strong written and verbal communication skills, including incident documentation and postmortem authoring. • Experience working in start-up, scale-up, or other fast-moving engineering environments. • Ability to complete a background check to company standard. • Willingness to participate in separate rotating weekday and weekend on-call schedules.

🏖️ Benefits

• Benefits of a dynamic and growing organization • A challenging, hands-on role that will have real impact on the business • Competitive compensation • Comprehensive benefits package • 401(k) Savings Plan • Medical/Dental/Vision Benefits • Health Savings Account (HSA) and Flexible Spending Account (FSA) • Unlimited Paid Time-off • 11 Paid Holidays • Paid Parental Leave • Company Paid Short-term Disability (STD) • Company Paid Long-term Disability (LTD) • Company Paid Life Insurance

Apply Now

Similar Jobs

🔥 1 hour ago

Synapticure Inc.

11 - 50

🏥 Healthcare

📡 Telecommunications

⚕️ Healthcare Insurance

Senior DevOps and Security Engineer securing AWS, Kubernetes, and software delivery environments. Supporting Synapticure’s virtual neurodegenerative disease care and life sciences research platform.

🔥 1 hour ago

Wursta

51 - 200

đź’Ľ Consulting

🏥 Healthcare

📦 Logistics

Digital Workplace Deployment Engineer leading Google Workspace migrations and cloud deployment services. Delivering digital transformation, managed services, cybersecurity, and AI solutions at Wursta.

🔥 1 hour ago

Humana

10,000+ employees

🏥 Healthcare

🛡️ Insurance

⚕️ Healthcare Insurance

Senior DevOps Engineer advancing AI/ML DevOps maturity, cloud automation, and CI/CD pipelines at Humana. Improving security, releases, infrastructure provisioning, and software quality.

🔥 3 hours ago

Baxter International Inc.

10,000+ employees

🏥 Healthcare

🏭 Manufacturing

⚕️ Healthcare Insurance

Senior Site Reliability Engineer operating Baxter’s Azure-based digital health platforms. Ensuring 24x7 reliability, compliance, monitoring, disaster recovery, and business continuity.

🔥 3 hours ago

Bitsight

501 - 1000

Senior DevOps Engineer operating and automating Bitsight’s SaaS cloud infrastructure. Leveraging AI, Kubernetes, Terraform, and observability to improve secure, reliable cyber risk software delivery.