
10,000+ employees
Founded 2004
đą Media
Media ⢠Entertainment
NBCUniversal is a leading global media and entertainment company known for creating and distributing content across a variety of platforms. With over 100 years of experience, it is a part of Comcast and encompasses brands like Peacock, NBC Sports, and many others to educate, entertain, and empower audiences around the world. The company is involved in television broadcasting, film production, and theme parks, and is also recognized for its initiatives in technology and corporate social responsibility. NBCUniversal is committed to innovation and social impact, making it a vibrant workplace for media and tech professionals.
đ Yesterday
đ¨đŚ Canada â Remote
đľ $180k - $230k / year
â° Full Time
đ´ Lead
â DevOps & Site Reliability Engineer (SRE)
đť Ghost score 0%
Improve your chances of getting an interview by checking your resume score before you apply.

10,000+ employees
Founded 2004
đą Media
Media ⢠Entertainment
NBCUniversal is a leading global media and entertainment company known for creating and distributing content across a variety of platforms. With over 100 years of experience, it is a part of Comcast and encompasses brands like Peacock, NBC Sports, and many others to educate, entertain, and empower audiences around the world. The company is involved in television broadcasting, film production, and theme parks, and is also recognized for its initiatives in technology and corporate social responsibility. NBCUniversal is committed to innovation and social impact, making it a vibrant workplace for media and tech professionals.
⢠Architect and evolve the Kubernetes-native platform powering NBC broadcast production environments ⢠Design platform infrastructure that automates provisioning, lifecycle management, and cloud infrastructure delivery at enterprise scale ⢠Model broadcast infrastructure as custom resources using Crossplane compositions and custom Go functions ⢠Lead provisioning automation across multi-account AWS environments and on-premises control rooms ⢠Design, build, and maintain Kubernetes operators, controllers, and internal platform APIs in Go ⢠Develop custom Crossplane providers integrating enterprise platforms such as NRCS, Venafi, and Infoblox ⢠Manage resource lifecycles and approval workflows ⢠Lead cloud networking, DNS strategies, cross-account connectivity, VPC topology, and dynamic network routing ⢠Partner with broadcast systems engineers, system integrators, and external vendors ⢠Automate bare-metal compute configurations with Puppet and integrate proprietary vendor solutions ⢠Write RFCs, drive architectural decisions, mentor engineers, and establish CI/CD and testing strategies ⢠Own authorization models, hierarchical RBAC, resource identifiers, and identity integrations ⢠Drive GitOps continuous delivery using Flux, Kustomize, and Helm ⢠Manage configuration-as-code for compute fleets using Puppet ⢠Design observability and alerting stacks ⢠Oversee remote desktop/VDI connectivity, authentication, credential management, and gateway routing ⢠Contribute upstream to open-source projects and improve cloud-native solutions
⢠10+ years of experience designing, building, and operating production infrastructure and cloud-native platforms at enterprise scale ⢠Strong proficiency in Go, including systems-level programming and API servers ⢠Deep experience building Kubernetes controllers/operators using controller-runtime and kubebuilder ⢠Expert-level knowledge of Kubernetes, including CRD/XRD generation, operators, informers, admission webhooks, and RBAC ⢠Deep production experience with Crossplane, including composite resources, composition functions, and custom Crossplane providers in Go ⢠Extensive production experience with AWS multi-account architectures, cross-account networking, and identity federation ⢠Experience with EKS, EC2, VPC, IAM, STS, SSM, Secrets Manager, Route 53, and S3 ⢠Production experience with GitOps tooling, specifically Flux or ArgoCD ⢠Hands-on Puppet experience, including module development, PuppetDB, Hiera, and r10k ⢠Experience designing REST APIs with middleware patterns and OAuth/JWT authentication ⢠Knowledge of information security, IAM trust chains, least-privilege policies, JWT lifecycles, and secrets abstraction ⢠Experience designing observability platforms using Grafana, Prometheus/Mimir, Loki, OpenTelemetry, Alloy, or Prometheus Node Exporter ⢠Working knowledge of PostgreSQL, SQLite, or similar relational databases, including schema design, migrations, and query optimization ⢠Ability to present architectural decisions, engage with vendors, and write technical documentation ⢠Familiarity with broadcast/media production workflows and live production constraints preferred ⢠Experience with Crossplane function SDK and Kubernetes disaster recovery preferred ⢠Familiarity with VDI solutions, machine identity workflows, and PKI certificate management preferred ⢠Experience with hybrid DNS, software-defined networking, Envoy Gateway, or Gateway API preferred ⢠Familiarity with k6, KUTTL, SOPS, Air, kind, or colima preferred ⢠Ability to script in Bash and PowerShell ⢠Active open-source contributions, particularly in the CNCF/Kubernetes ecosystem, preferred
⢠Medical insurance ⢠Dental insurance ⢠Vision insurance ⢠401(k) ⢠Paid leave ⢠Tuition reimbursement ⢠Discounts and perks ⢠Bonus eligibility
Apply Nowđ August 27
Staff DevOps Engineer owning reliable, scalable infrastructure for Nexxaâs AI systems. Supporting machine learning workloads across heavy-industry operations.
đ August 18
Staff Site Reliability Engineer strengthening AWS and Kubernetes resilience for Caseware, a fintech company building audit and accounting software. Driving secure delivery, observability, and incident management.
đ August 5
Staff SRE leading BeyondTrustâs Password Safe identity-security platform across cloud and on-premises infrastructure. Driving GitOps, CI/CD, observability, resilience, and reliability strategy.
đ¨đŚ Canada â Remote
đ° Private Equity Round on 2021-05
â° Full Time
đ´ Lead
â DevOps & Site Reliability Engineer (SRE)
đ July 16
Director of SRE leading global infrastructure and reliability teams for Blackpoint Cyber's cyber defense services. Focusing on cost efficiency, automation, and team leadership.
đ¨đŚ Canada â Remote
đľ CA$167k - CA$213k / year
đ° $190M Series C on 2023-06
â° Full Time
đ´ Lead
â DevOps & Site Reliability Engineer (SRE)
đ July 1
AI DevOps & Reliability Engineer at Branch, focusing on software delivery and operational standards, enhancing DevOps with AI tools for reliability and efficiency.
đ¨đŚ Canada â Remote
đľ $123k - $160k / year
đ° $300M Series F - Branch on 2022-02
â° Full Time
đ Senior
đ´ Lead
â DevOps & Site Reliability Engineer (SRE)