Search Remote Jobs

Distinguished Engineer, Core DevOps

🕒 Yesterday

🌐 United States, Canada, +1 more countries – Remote

infoinfo

💵 $250k - $349k / year

⏰ Full Time

🟠 Senior

🔴 Lead

⛑ DevOps & Site Reliability Engineer (SRE)

👻 Ghost score 0%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of GitLab

GitLab

1001 - 5000 employees

Founded 2014

💼 Consulting

📣 Marketing

🤖 Artificial Intelligence

💰 Secondary Market on 2020-11

Consulting • Marketing • Artificial Intelligence

GitLab is the most comprehensive AI-powered DevSecOps platform, offering tools for automated software delivery, security, and compliance throughout the software development lifecycle. It provides solutions across areas such as AI-assisted development, continuous integration/continuous deployment (CI/CD), source code management, and vulnerability management. GitLab aims to simplify and accelerate software delivery by uniting development, security, and operations on a unified platform. It is particularly recognized for its AI code assistants and has been named a leader in the Gartner Magic Quadrant™ for DevOps Platforms, making it a preferred choice for many enterprises.

📋 Description

• Set and continuously refine the go-forward technical direction for CI, CD, Plan, and source code experiences. • Partner directly with the VP of Engineering on planning, technical investment, and engineering priorities. • Translate company-wide productivity, quality, and AI-native engineering goals into technical direction and iterative roadmaps for Core DevOps. • Define measurable quality standards for availability, p95 and p99 latency, pipeline success rate, test reliability, data correctness, and defect escape rate. • Identify systemic technical risks and drive prioritized plans to retire them. • Lead complex design decisions through code, prototypes, and proofs of concept. • Drive convergence on shared patterns, libraries, and paved paths. • Design incremental migration, dual-run, rollback, and deprecation strategies for long-lived systems. • Adopt AI platform capabilities into Core DevOps and feed requirements back to the AI organization. • Serve as escalation point for complex or contested technical decisions. • Review critical-path designs and merge requests, using reviews to teach. • Ensure designs work across GitLab.com, GitLab Dedicated, and Self-Managed deployments while respecting multi-tenant, compliance, and data-governance boundaries. • Partner with Infrastructure, Security, and SRE on observability, debuggability, and graceful failure modes. • Communicate architectural constraints and opportunities to Product through roadmap terms. • Write design documents, architecture narratives, and decision records. • Grow Principal and Staff Engineers and contribute to the senior technical hiring bar. • Participate in the Incident Management on-call rotation and complete Interview Training for technical interviewing.

🎯 Requirements

• 10+ years of software engineering experience, including 4+ years in a Staff, Principal, or equivalent senior technical leadership role. • Deep expertise in AI and ML systems, including large language models, agentic frameworks, and autonomous workflow design at production scale. • Proven track record of leading hands-on technical experimentation, including defining evaluation frameworks, running benchmarks, and translating findings into scalable architecture decisions. • Strong background in scalable, multi-tenant distributed systems, including service decomposition, fault tolerance, observability, and operational resilience. • Experience designing and implementing human-in-the-loop controls, safety guardrails, and responsible AI practices for production systems. • Experience mentoring senior engineers and influencing technical direction across multiple teams or divisions without direct authority. • Ability to work effectively in a fully remote, globally distributed organization with excellent written and asynchronous communication skills.

🏖️ Benefits

• Benefits to support your health, finances, and well-being • Flexible Paid Time Off • Team Member Resource Groups • Equity Compensation & Employee Stock Purchase Plan • Growth and Development Fund • Parental Leave

Apply Now

Similar Jobs

🕒 Yesterday

Vultr

201 - 500

🤖 Artificial Intelligence

🤝 B2B

🔧 Hardware

Senior SRE maintaining MySQL and PostgreSQL reliability for Vultr’s global cloud infrastructure. Owning monitoring, disaster recovery, incident response, security compliance, and automation.

🕒 Yesterday

Horizon3.ai

51 - 200

🔒 Cybersecurity

🤖 Artificial Intelligence

☁️ SaaS

Staff SRE owning reliability strategy, observability, and incident readiness for Horizon3’s autonomous cybersecurity platform. Leading cross-functional initiatives across production infrastructure and services.

🕒 Yesterday

Cisco

10,000+ employees

🔧 Hardware

🔐 Security

🏢 Enterprise

Site Reliability Engineer supporting Cisco-owned Splunk’s FedRAMP cloud platform. Testing features, automating infrastructure, and leading customer incidents on overnight remote shifts.

🕒 Yesterday

CACI International Inc

10,000+ employees

🎖️ Defense

🏛️ Government

🔒 Cybersecurity

Operating System Deployment Engineer managing secure Windows and Windows Server images for CACI’s DoD enterprise IT services. Automating deployments and supporting physical and virtual infrastructure across 187 bases.

🇺🇸 United States – Remote

💵 $75.2k - $158.1k / year

🔥 Funding within the last year

💰 $500M Post-IPO Debt on 2026-02

⏰ Full Time

🟠 Senior

🔴 Lead

⛑ DevOps & Site Reliability Engineer (SRE)

🕒 Yesterday

PingWind Inc. (SDVOSB)

51 - 200

💼 Consulting

📦 Logistics

🏥 Healthcare

DevSecOps Engineer securing cloud-native software delivery for PingWind, a federal government services provider. Building CI/CD security, automating controls, and supporting vulnerability remediation.