
10,000+ employees
Founded 1915
🏗️ Construction
🏥 Healthcare
📦 Logistics
Construction • Healthcare • Logistics
Carrier is a global leader in building and cold chain solutions, dedicated to innovation in creating healthy, safe, sustainable, and intelligent environments. The company is involved in HVAC and refrigeration systems and focuses on promoting the health and safety of indoor spaces and preserving the global supply of food and medicine through its advanced technologies. Carrier is also committed to addressing climate change and collaborates with partners to drive sustainability and energy efficiency in infrastructure. By exploring smart building solutions and energy resiliency, Carrier is making strides towards a net zero future.
🔥 0 minutes ago
🐊 Florida, New Mexico, +2 more states – Remote
💵 $96k - $192k / year
⏰ Full Time
🔴 Lead
⛑ DevOps & Site Reliability Engineer (SRE)
🦅 H1B Visa Sponsor
👻 Ghost score 0%
Improve your chances of getting an interview by checking your resume score before you apply.

10,000+ employees
Founded 1915
🏗️ Construction
🏥 Healthcare
📦 Logistics
Construction • Healthcare • Logistics
Carrier is a global leader in building and cold chain solutions, dedicated to innovation in creating healthy, safe, sustainable, and intelligent environments. The company is involved in HVAC and refrigeration systems and focuses on promoting the health and safety of indoor spaces and preserving the global supply of food and medicine through its advanced technologies. Carrier is also committed to addressing climate change and collaborates with partners to drive sustainability and energy efficiency in infrastructure. By exploring smart building solutions and energy resiliency, Carrier is making strides towards a net zero future.
• Design, implement, and maintain highly available, scalable, and secure cloud infrastructure supporting Carrier's SaaS platforms • Define and drive SLOs, SLIs, and SLAs across critical services • Build self-service platform capabilities for efficient and consistent service deployment and operation • Develop reliability frameworks, standards, and operational best practices • Identify reliability bottlenecks and eliminate single points of failure • Develop automation to eliminate repetitive operational tasks and reduce manual intervention • Design and maintain Infrastructure as Code using Terraform, AWS CloudFormation, AWS CDK, or similar technologies • Build and enhance CI/CD pipelines • Implement automated remediation, self-healing capabilities, and operational workflows • Design and enhance observability solutions using metrics, logs, traces, and distributed monitoring • Develop dashboards, monitoring standards, alerting frameworks, and reliability reporting • Lead critical production incident response, root cause analysis, and long-term corrective actions • Drive post-incident reviews focused on systemic improvements • Improve platform performance, availability, scalability, disaster recovery, and business continuity • Validate backup, restoration, and recovery processes through regular testing • Collaborate on resilient architectures meeting reliability and compliance objectives • Support operational readiness reviews for new services and platform capabilities • Mentor engineers and provide technical leadership • Partner with development teams throughout the software development lifecycle • Promote automation, ownership, continuous learning, and operational excellence • Stay current on cloud infrastructure, platform engineering, AI-assisted operations, and SRE practices • Participate in a shared on-call rotation • Drive initiatives reducing operational toil through automation, self-healing systems, and platform improvements
• Bachelor's degree in Computer Science, Software Engineering, Information Technology, or a technical field with 7+ years of experience in Site Reliability Engineering, Platform Engineering, DevOps, or Cloud Engineering, OR Master’s degree in one of these fields with 5+ years of experience in these areas • 3+ years of hands-on experience designing and supporting large-scale cloud environments • 3+ years of experience with Infrastructure as Code, such as Terraform, CloudFormation, or AWS CDK • 3+ years of experience building and maintaining CI/CD pipelines and deployment automation • 3+ years of extensive experience with Amazon Web Services, including EKS, EC2, VPC, RDS, Lambda, CloudWatch, and IAM • 3+ years of experience implementing observability platforms using Prometheus, Grafana, Open Telemetry, Datadog, Splunk, or New Relic • Strong understanding of modern SRE principles, including reliability engineering, observability, toil reduction, incident management, and error budgets • Strong scripting or programming experience using Python, Go, PowerShell, Bash, or similar languages • Strong knowledge of networking, security, systems architecture, and distributed systems concepts • Proven experience leading production incident response and root cause analysis activities • Proven experience with cloud-native platforms and container technologies, including Kubernetes and Docker • Experience operating SaaS platforms serving large-scale customer environments • Experience supporting compliance frameworks such as SOC 2, ISO 27001, NIST, or similar standards • Experience implementing AI-assisted engineering solutions to improve operational efficiency, troubleshooting, automation, and service reliability • Knowledge of platform engineering concepts including Internal Developer Platforms, developer self-service capabilities, and engineering enablement practices • AWS certifications or other cloud certifications • Excellent communication, leadership, and collaboration skills
• Short-term cash incentives, subject to plan requirements • Medical, Dental, Vision • Wellness incentives • Retirement Benefits • Paid vacation days, up to 15 days • Paid sick days, up to 5 days • Paid personal leave, up to 5 days • Paid holidays, up to 13 days • Birth and adoption leave • Parental leave • Family and medical leave • Bereavement leave • Jury duty leave • Military leave • Purchased vacation • Short-term and long-term disability • Life Insurance and Accidental Death and Dismemberment • Health Savings Account • Health Care Spending Account • Dependent Care Spending Account • Tuition Assistance
Apply Now🔥 13 hours ago
SRE Delivery Manager leading SRE and Delivery teams for Ninety’s EOS business-management software. Improving AWS infrastructure, deployments, observability, security, and incident response.
🇺🇸 United States – Remote
💵 $200k - $220k / year
💰 $35M Series B - Ninety on 2023-11
⏰ Full Time
🟠 Senior
🔴 Lead
⛑ DevOps & Site Reliability Engineer (SRE)
🕒 Yesterday
Principal Deployment Engineer serving as senior engineering contact for sensitive customers. Leading IT deployments, testing, training, and customer support for AIS cyber and information security operations.
🇺🇸 United States – Remote
💵 $126k - $180k / year
💰 Venture Round on 2015-06
⏰ Full Time
🔴 Lead
⛑ DevOps & Site Reliability Engineer (SRE)
🕒 Yesterday
Principal SRE scaling Kubernetes infrastructure and Golang services for Blue River’s autonomous robotics platforms. Driving reliability, observability, security, and platform adoption across engineering teams.
🇺🇸 United States – Remote
💵 $174k - $305k / year
⏰ Full Time
🔴 Lead
⛑ DevOps & Site Reliability Engineer (SRE)
🦅 H1B Visa Sponsor
🕒 Yesterday
Staff Site Reliability Engineer operating AWS and Kubernetes infrastructure for Butterfly Network’s medical ultrasound platform. Improving observability, reliability, and incident response across clinical workflows.
🇺🇸 United States – Remote
💵 $190k - $210k / year
💰 $75.6M Post-IPO Equity - Butterfly Network on 2025-01
⏰ Full Time
🔴 Lead
⛑ DevOps & Site Reliability Engineer (SRE)
🕒 2 days ago
DevSecOps Engineer building and deploying secure cloud-based IAM systems for federal government clients. Maintaining highly available architectures, automated delivery, compliance, and infrastructure upgrades.
🇺🇸 United States – Remote
💵 $93k - $128k / year
⏰ Full Time
🟠 Senior
🔴 Lead
⛑ DevOps & Site Reliability Engineer (SRE)