Search Remote Jobs

Principal DevOps Engineer

🕒 Yesterday

🇨🇦 Canada – Remote

💵 $180k - $230k / year

⏰ Full Time

🔴 Lead

⛑ DevOps & Site Reliability Engineer (SRE)

👻 Ghost score 0%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of NBCUniversal

NBCUniversal

10,000+ employees

Founded 2004

📱 Media

Media • Entertainment

NBCUniversal is a leading global media and entertainment company known for creating and distributing content across a variety of platforms. With over 100 years of experience, it is a part of Comcast and encompasses brands like Peacock, NBC Sports, and many others to educate, entertain, and empower audiences around the world. The company is involved in television broadcasting, film production, and theme parks, and is also recognized for its initiatives in technology and corporate social responsibility. NBCUniversal is committed to innovation and social impact, making it a vibrant workplace for media and tech professionals.

📋 Description

• Architect and evolve the Kubernetes-native platform powering NBC broadcast production environments • Design platform infrastructure that automates provisioning, lifecycle management, and cloud infrastructure delivery at enterprise scale • Model broadcast infrastructure as custom resources using Crossplane compositions and custom Go functions • Lead provisioning automation across multi-account AWS environments and on-premises control rooms • Design, build, and maintain Kubernetes operators, controllers, and internal platform APIs in Go • Develop custom Crossplane providers integrating enterprise platforms such as NRCS, Venafi, and Infoblox • Manage resource lifecycles and approval workflows • Lead cloud networking, DNS strategies, cross-account connectivity, VPC topology, and dynamic network routing • Partner with broadcast systems engineers, system integrators, and external vendors • Automate bare-metal compute configurations with Puppet and integrate proprietary vendor solutions • Write RFCs, drive architectural decisions, mentor engineers, and establish CI/CD and testing strategies • Own authorization models, hierarchical RBAC, resource identifiers, and identity integrations • Drive GitOps continuous delivery using Flux, Kustomize, and Helm • Manage configuration-as-code for compute fleets using Puppet • Design observability and alerting stacks • Oversee remote desktop/VDI connectivity, authentication, credential management, and gateway routing • Contribute upstream to open-source projects and improve cloud-native solutions

🎯 Requirements

• 10+ years of experience designing, building, and operating production infrastructure and cloud-native platforms at enterprise scale • Strong proficiency in Go, including systems-level programming and API servers • Deep experience building Kubernetes controllers/operators using controller-runtime and kubebuilder • Expert-level knowledge of Kubernetes, including CRD/XRD generation, operators, informers, admission webhooks, and RBAC • Deep production experience with Crossplane, including composite resources, composition functions, and custom Crossplane providers in Go • Extensive production experience with AWS multi-account architectures, cross-account networking, and identity federation • Experience with EKS, EC2, VPC, IAM, STS, SSM, Secrets Manager, Route 53, and S3 • Production experience with GitOps tooling, specifically Flux or ArgoCD • Hands-on Puppet experience, including module development, PuppetDB, Hiera, and r10k • Experience designing REST APIs with middleware patterns and OAuth/JWT authentication • Knowledge of information security, IAM trust chains, least-privilege policies, JWT lifecycles, and secrets abstraction • Experience designing observability platforms using Grafana, Prometheus/Mimir, Loki, OpenTelemetry, Alloy, or Prometheus Node Exporter • Working knowledge of PostgreSQL, SQLite, or similar relational databases, including schema design, migrations, and query optimization • Ability to present architectural decisions, engage with vendors, and write technical documentation • Familiarity with broadcast/media production workflows and live production constraints preferred • Experience with Crossplane function SDK and Kubernetes disaster recovery preferred • Familiarity with VDI solutions, machine identity workflows, and PKI certificate management preferred • Experience with hybrid DNS, software-defined networking, Envoy Gateway, or Gateway API preferred • Familiarity with k6, KUTTL, SOPS, Air, kind, or colima preferred • Ability to script in Bash and PowerShell • Active open-source contributions, particularly in the CNCF/Kubernetes ecosystem, preferred

🏖️ Benefits

• Medical insurance • Dental insurance • Vision insurance • 401(k) • Paid leave • Tuition reimbursement • Discounts and perks • Bonus eligibility

Apply Now

Similar Jobs

🕒 August 27

Nexxa.ai

11 - 50

🏗️ Construction

🏭 Manufacturing

📦 Logistics

Staff DevOps Engineer owning reliable, scalable infrastructure for Nexxa’s AI systems. Supporting machine learning workloads across heavy-industry operations.

🇨🇦 Canada – Remote

⏰ Full Time

🔴 Lead

⛑ DevOps & Site Reliability Engineer (SRE)

🕒 August 18

Caseware

201 - 500

💸 Finance

🏢 Enterprise

☁️ SaaS

Staff Site Reliability Engineer strengthening AWS and Kubernetes resilience for Caseware, a fintech company building audit and accounting software. Driving secure delivery, observability, and incident management.

🇨🇦 Canada – Remote

⏰ Full Time

🔴 Lead

⛑ DevOps & Site Reliability Engineer (SRE)

🕒 August 5

BeyondTrust

1001 - 5000

🔒 Cybersecurity

Staff SRE leading BeyondTrust’s Password Safe identity-security platform across cloud and on-premises infrastructure. Driving GitOps, CI/CD, observability, resilience, and reliability strategy.

🇨🇦 Canada – Remote

💰 Private Equity Round on 2021-05

⏰ Full Time

🔴 Lead

⛑ DevOps & Site Reliability Engineer (SRE)

🕒 July 16

Blackpoint Cyber

51 - 200

💼 Consulting

🎖️ Defense

🔒 Cybersecurity

Director of SRE leading global infrastructure and reliability teams for Blackpoint Cyber's cyber defense services. Focusing on cost efficiency, automation, and team leadership.

🇨🇦 Canada – Remote

💵 CA$167k - CA$213k / year

💰 $190M Series C on 2023-06

⏰ Full Time

🔴 Lead

⛑ DevOps & Site Reliability Engineer (SRE)

🕒 July 1

Branch

201 - 500

📣 Marketing

☁️ SaaS

🤝 B2B

AI DevOps & Reliability Engineer at Branch, focusing on software delivery and operational standards, enhancing DevOps with AI tools for reliability and efficiency.

🇨🇦 Canada – Remote

💵 $123k - $160k / year

💰 $300M Series F - Branch on 2022-02

⏰ Full Time

🟠 Senior

🔴 Lead

⛑ DevOps & Site Reliability Engineer (SRE)