Engineering Manager, Site Reliability

🔥 0 minutes ago

🌐 United States, Canada – Remote

infoinfo

💵 $178k - $226k / year

⏰ Full Time

🟠 Senior

🔴 Lead

⛑ DevOps & Site Reliability Engineer (SRE)

🦅 H1B Visa Sponsor

infoinfo

👻 Ghost score 0%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Instacart

Instacart

1001 - 5000 employees

Founded 2012

🍽️ Food & Beverage

📦 Logistics

🛍️ eCommerce

💰 $232M Venture Round on 2021-11

Food & Beverage • Logistics • eCommerce

Instacart is a company that offers a flexible approach to work while transforming the grocery industry. It provides an essential service by delivering groceries and household goods to customers' doors in as little as 30 minutes. Instacart offers safe and flexible earning opportunities to personal shoppers and tackles challenges such as rerouting deliveries during snowstorms and connecting customers with coupons and deals. It aims to be the operating system for the grocery industry, thus helping customers save time for other activities. Instacart emphasizes diversity, equity, and belonging in its work culture.

📋 Description

• Lead, mentor, and develop a team of Site Reliability Engineers • Establish team goals, provide actionable feedback, and support career growth and professional development • Set technical direction and priorities for reliability, scalability, availability, performance, and operational readiness • Partner with engineering, product, security, infrastructure, and other cross-functional teams • Define reliability standards, influence system design, and deliver initiatives improving customer and developer experience • Drive incident management, incident response, post-incident learning, service-level objectives, capacity planning, observability, and continuous risk reduction • Promote automation and self-service tooling to reduce operational toil and improve deployment confidence • Balance near-term operational needs with long-term investments • Communicate reliability risks, project status, tradeoffs, and decisions to technical and non-technical stakeholders • Make decisions under ambiguity, respond calmly during incidents, and help teams learn from failure

🎯 Requirements

• Bachelor’s degree in Computer Science, Engineering, or a related technical field, or equivalent practical experience • Seven or more years of experience in software engineering, infrastructure engineering, Site Reliability Engineering, or a related field • Two or more years of experience managing, mentoring, or leading engineering teams • Professional experience with cloud infrastructure, distributed systems, networking, containers, orchestration platforms, or related production technologies • Experience leading or participating in production incident response, post-incident reviews, reliability improvement initiatives, and operational readiness practices • Experience communicating technical risks, priorities, and tradeoffs to engineering leaders and cross-functional stakeholders • Preferred: experience leading Site Reliability Engineering, platform engineering, infrastructure engineering, or developer productivity teams • Preferred: experience operating highly available services at significant scale and improving service-level objectives, observability, capacity, or disaster recovery capabilities • Preferred: experience with infrastructure as code, continuous delivery, monitoring, logging, tracing, and automated remediation • Preferred: experience building or evolving reliability programs across multiple engineering teams • Preferred: demonstrated ability to create alignment across teams, navigate ambiguity, and turn complex technical challenges into clear, achievable plans • Preferred: leadership approach grounded in empathy, transparency, direct communication, collaboration, and inclusive team development

🏖️ Benefits

• Flexibility to work from home, an office, or a coffee shop • Regular in-person events • New hire equity grant • Annual refresh equity grants • Market-competitive compensation and benefits

Apply Now

Similar Jobs

🔥 1 hour ago

Cisco

10,000+ employees

🔧 Hardware

🔐 Security

🏢 Enterprise

Cisco SRE technical leader migrating specialized Kubernetes workloads from AWS to internal platforms. Operating globally distributed services and improving Kubernetes reliability, networking, and scalability.

🔥 1 hour ago

Databento

11 - 50

💼 Consulting

📣 Marketing

📦 Logistics

Site Reliability Engineer maintaining reliability, performance, and observability for Databento’s next-generation financial market-data platform. Improving deployments, incident response, and backend infrastructure.

🔥 1 hour ago

Peraton

10,000+ employees

💼 Consulting

🏥 Healthcare

📦 Logistics

Senior SRE building Python, AWS, and Terraform reliability solutions for Peraton’s national security missions. Improving cloud platform reliability through automation, observability, and incident management.

🔥 9 hours ago

The SSI Group

11 - 50

🎯 Recruiter

👥 HR Tech

🤝 B2B

DevOps Process Analyst optimizing development and operations workflows for SSI Group, a healthcare software company. Improving processes, reporting, risk management, and cross-functional delivery.

🔥 11 hours ago

CentralReach

201 - 500

🏥 Healthcare

📚 Education

Site Reliability Engineer operating AWS and Snowflake infrastructure for CentralReach’s autism and IDD care software. Improving reliability, connectivity, deployments, and incident response across data platforms.

🇺🇸 United States – Remote

💵 $135k - $160k / year

💰 Private equity on 2018-03

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)