Principal Site Reliability Engineer, Platform Engineering

Job not on LinkedIn

🔥 16 hours ago

🌐 United States, Canada, +1 more countries – Remote

infoinfo

💵 $223.2k - $380.4k / year

⏰ Full Time

🔴 Lead

⛑ DevOps & Site Reliability Engineer (SRE)

👻 Ghost score 0%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of GitLab

GitLab

1001 - 5000 employees

Founded 2014

💼 Consulting

📣 Marketing

🤖 Artificial Intelligence

💰 Secondary Market on 2020-11

Consulting • Marketing • Artificial Intelligence

GitLab is the most comprehensive AI-powered DevSecOps platform, offering tools for automated software delivery, security, and compliance throughout the software development lifecycle. It provides solutions across areas such as AI-assisted development, continuous integration/continuous deployment (CI/CD), source code management, and vulnerability management. GitLab aims to simplify and accelerate software delivery by uniting development, security, and operations on a unified platform. It is particularly recognized for its AI code assistants and has been named a leader in the Gartner Magic Quadrant™ for DevOps Platforms, making it a preferred choice for many enterprises.

📋 Description

• Set technical direction for GitLab Dedicated, shaping architecture and platform strategy as we scale a growing fleet of isolated, single-tenant environments • Lead platform transformations across resilience, failover, tenant orchestration, change management, self-service tooling, and platform integrations • Drive scalable, modular architecture aligned with GitLab’s broader Cells strategy while preserving security, isolation, and compliance requirements • Strengthen service ownership and operational maturity, helping engineering teams build, operate, and improve the production systems they own • Identify and address systemic reliability and scalability risks using production signals, incident patterns, and architectural insight • Establish reusable platform patterns and automation that reduce operational toil and allow Dedicated to scale efficiently • Lead complex technical decisions across teams, balancing reliability, security, cost, maintainability, and customer needs • Advance engineering excellence through architectural leadership, mentorship, and influence with senior engineers and engineering leaders

🎯 Requirements

• Deep expertise in Site Reliability, Platform, Infrastructure, or Backend Engineering, with experience designing and operating large-scale production systems • Hands-on experience with cloud infrastructure, automation, observability, infrastructure as code, and modern production engineering practices • Strong software engineering fundamentals, with experience building production systems or infrastructure tooling in languages such as Go, Ruby, Python, or similar • Strong distributed systems and systems-design expertise, with sound judgment around reliability, failure isolation, scalability, and operational complexity • A track record of technical leadership across multiple teams, setting direction and driving complex initiatives through influence • Experience leading significant platform or infrastructure transformations, including modernization, modularization, or scaling systems through major growth • Experience leading changes that improve how engineering teams own and operate production systems, strengthening reliability, operational readiness, and accountability at scale • Exceptional technical communication and influence, with the ability to build alignment, mentor senior engineers, and guide complex architectural decisions

🏖️ Benefits

• Benefits to support your health, finances, and well-being • Flexible Paid Time Off • Team Member Resource Groups • Equity Compensation & Employee Stock Purchase Plan • Growth and Development Fund • Parental Leave

Apply Now

Similar Jobs

🕒 Yesterday

Sparibis

11 - 50

💼 Consulting

🔒 Cybersecurity

🏢 Enterprise

DevSecOps Engineer building secure AWS cloud infrastructure and CI/CD pipelines for Sparibis, a professional solutions firm. Automating deployments, security controls, and AI-assisted development workflows.

🕒 Yesterday

Velera

1001 - 5000

💼 Consulting

📣 Marketing

💳 Fintech

Technology Delivery Manager leading Agile software delivery, Azure DevOps, and team development for Velera, a payments fintech serving credit unions. Driving quality, architecture, and continuous improvement.

🕒 Yesterday

NVIDIA

10,000+ employees

🏥 Healthcare

🏭 Manufacturing

🤖 Artificial Intelligence

Engineering Manager leading NVIDIA’s EDA infrastructure reliability team. Owning operational platforms, incident readiness, automation, and engineering delivery for chip development systems.

🕒 Yesterday

DuckDuckGo

51 - 200

💼 Consulting

📣 Marketing

🔒 Cybersecurity

Director of Site Reliability Engineering leading scalable infrastructure and reliability initiatives. DuckDuckGo protects users online through private browsing, search, subscription, and AI products.

🕒 Yesterday

ExpertVoice

51 - 200

📣 Marketing

🛒 Retail

🛍️ eCommerce

Site Reliability Engineer IV setting technical direction for ExpertVoice’s platform serving leading consumer brands. Owning reliability, automation, CI/CD, architecture, and incident response at scale.