Principal DevOps Engineer

🕒 Julho 1

🗽 New York – Remoto

info

💵 $180.000 - $230.000 / ano

⏰ Tempo Integral

🔴 Especialista

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🦅 Patrocina Visto H1B

info

🗣️🇺🇸🇬🇧 Inglês obrigatório

Candidatar-se
Encontrar Vagas Remotas Similares

📊 Verifique sua pontuação de currículo para esta vaga

Melhore suas chances de conseguir uma entrevista verificando sua pontuação de currículo antes de se candidatar.

Logo of NBCUniversal

NBCUniversal

10.000+ funcionários

Fundada em 2004

📱 Mídia

Media • Entertainment

Criamos conteúdo de classe mundial, que distribuímos em nosso portfólio de filmes, televisão e streaming, e ganham vida por meio de nossos parques temáticos e experiências para o consumidor. Detemos e operamos marcas líderes de entretenimento e notícias, incluindo NBC, NBC News, MSNBC, CNBC, NBC Sports, Telemundo, NBC Local Stations, Bravo, USA Network e Peacock, nosso serviço de streaming premium com anúncios. Produzimos e distribuímos entretenimento audiovisual e programação de primeira linha por meio da Universal Filmed Entertainment Group e da Universal Studio Group, e contamos com parques temáticos e atrações de renome mundial por meio da Universal Destinations & Experiences. A NBCUniversal é uma subsidiária da Comcast Corporation.

Descrição

• Architect a Kubernetes-native platform that models broadcast infrastructure as custom resources. • Lead the technical strategy leveraging Crossplane compositions and custom Go functions to automate provisioning across multi-account AWS environments and on-prem control rooms. • Design, build, and maintain production-grade Kubernetes operators, controllers, and internal platform APIs in Go. • Actively develop custom Crossplane providers to deeply integrate external enterprise platforms (such as NRCS, Venafi, and Infoblox) into our control plane, managing resource lifecycles and approval workflows. • Lead the design of cloud networking, DNS strategies, and cross-account connectivity across hybrid environments, automating VPC topology and dynamic network routing. • Partner closely with broadcast systems engineers, system integrators, and external vendors to bridge the gap between broadcast hardware and automated infrastructure. • Write RFCs, drive architectural decisions, mentor engineers, and establish high-confidence CI/CD pipelines, testing strategies, and GitHub Actions automation. • Own the platform's authorization model, designing hierarchical RBAC systems, resource identifier schemes, and identity integrations that enforce fine-grained access control. • Drive GitOps-based continuous delivery (Flux, Kustomize, Helm) and manage configuration-as-code for compute fleets using Puppet. • Ensure deep operational visibility by designing comprehensive observability and alerting stacks. • Oversee the integration of remote desktop/VDI connectivity solutions, focusing on session authentication, credential management, and gateway routing.

🎯 Requisitos

• 10+ years of experience designing, building, and operating production infrastructure and cloud-native platforms at enterprise scale. • Strong proficiency in Go (systems-level programming, API servers). • Expert-level knowledge of the Kubernetes ecosystem, including CRD/XRD generation, operators, informers, admission webhooks, and RBAC. • Deep production experience with Crossplane, including composite resources, composition functions, and specifically developing custom Crossplane providers in Go to integrate external enterprise platforms. • Extensive production experience with AWS multi-account architectures, cross-account networking patterns, and identity federation. • Production experience with GitOps tooling, specifically Flux (HelmRelease, Kustomization) or ArgoCD for continuous delivery on Kubernetes. • Hands-on experience with Puppet, including module development, PuppetDB, Hiera, and r10k. • Experience designing REST APIs with middleware patterns and modern authentication (OAuth/JWT). • Keen eye for information security, including cross-account IAM trust chains, least-privilege policies, JWT token lifecycles, and secrets abstraction. • Strong background in designing telemetry platforms using Grafana, Prometheus/Mimir, Loki, OpenTelemetry, and metrics collection agents (Alloy, Prometheus Node Exporter). • Working knowledge of PostgreSQL, SQLite or similar relational databases, encompassing schema design, migrations, and query optimization. • Excellent problem-solving skills with a proven ability to present architectural decisions to executives, engage with vendors, and write clear technical documentation.

🏖️ Benefícios

• Health insurance • Dental insurance • Vision insurance • 401(k) • Paid leave • Tuition reimbursement • Variety of discounts and perks

Candidatar-se

Vagas Similares

🕒 Junho 30

Assured

11 - 50

🛡️ Seguros

☁️ SaaS

🤖 Inteligência Artificial

Staff Site Reliability Engineer optimizing database systems for tech-driven insurance provider. Leading design, automation, and performance initiatives for a modern claims processing platform.

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Junho 30

DraftKings Inc.

1001 - 5000

📣 Marketing

💼 Consultoria

📦 Logística

Principal Site Reliability Engineer shaping the Kubernetes platform and infrastructure strategy at DraftKings. Leading modernization and reliability initiatives across engineering teams.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $200.000 - $250.000 / ano

⏰ Tempo Integral

🔴 Especialista

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Junho 29

Convoso

201 - 500

💼 Consultoria

📣 Marketing

📦 Logística

Director of DevOps leading a team of engineers at Convoso, an AI-powered contact center platform. Responsible for developing and optimizing the platform and ensuring service reliability.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $220.000 - $260.000 / ano

⏰ Tempo Integral

🔴 Especialista

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Junho 29

FluidStack

11 - 50

🤖 Inteligência Artificial

Principal Operations Engineer overseeing critical operations in data centers for Fluidstack. Leading on-call escalation, root cause analysis, and operational excellence in real-time situations.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $150.000 - $250.000 / ano

⏰ Tempo Integral

🔴 Especialista

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Junho 24

Redox

201 - 500

🏥 Saúde

⚕️ Seguro de Saúde

☁️ SaaS

DevSecOps Engineer ensuring secure software development at Redox, enhancing healthcare data exchange. Collaborating with platform engineers to implement security best practices across the AWS/EKS infrastructure.

🗣️🇺🇸🇬🇧 Inglês obrigatório