Customer Reliability Engineer

🕒 6 dias atrás

🗣️🇺🇸🇬🇧 Inglês obrigatório

Candidatar-se
Encontrar Vagas Remotas Similares

📊 Verifique sua pontuação de currículo para esta vaga

Melhore suas chances de conseguir uma entrevista verificando sua pontuação de currículo antes de se candidatar.

Logo of Cisco

Cisco

10.000+ funcionários

Fundada em 1984

🔧 Hardware

🔐 Segurança

🏢 Corporativo

Hardware • Security • Enterprise

A Cisco é uma empresa multinacional de tecnologia que fornece hardware de rede, software e serviços para empresas, provedores de serviços e governos. A empresa desenvolve roteadores, switches, transceptores ópticos, silício programável e plataformas de computação de borda, além de oferecer segurança, colaboração (Webex), observabilidade e software habilitado por IA e serviços de suporte para ajudar organizações a projetar, operar e proteger redes e data centers em larga escala. A Cisco também oferece serviços profissionais, treinamentos e soluções gerenciadas em nuvem para apoiar a transformação digital e infraestrutura preparada para IA.

Descrição

• Own Cisco Hypershield cases escalated from Cisco TAC through resolution, engaging customers directly as needed • Diagnose complex production failures across the N9300 Smart Switch fabric and on-premises Kubernetes controller • Localize faults across switching and forwarding, security services and enforcement, and control-plane layers • Understand customer architectures and configurations and diagnose failures in unfamiliar production environments • Reproduce customer failures, partner with engineering to drive fixes, and own the fix back to the customer • Convert individual cases into systemic improvements such as runbooks, diagnostics, knowledge-base content, and product feedback • Build the team’s proactive view of customer health through monitoring, tooling, and reliability practices

🎯 Requisitos

• Bachelor’s degree plus 8 years of experience, Master’s degree plus 6 years, or equivalent industry experience • Experience supporting enterprise customers in an escalation capacity • Experience diagnosing and resolving complex production incidents under SLA pressure in unfamiliar environments • Experience operating and troubleshooting Cisco Nexus / NX-OS in production; equivalent depth on another major vendor accepted • Experience localizing failures across layered data-center architectures spanning switching/forwarding, services/enforcement, and control-plane domains • Linux operations experience at the command line, including production troubleshooting • Working exposure to containers or Kubernetes • Experience with packet capture and flow-telemetry analysis, including NetFlow/IPFIX • Working knowledge of enterprise virtualization and troubleshooting VM-based appliance deployments; vSphere is the current deployment target • Operational Kubernetes and Helm proficiency, including TLS certificates, service-account authentication, API-server connectivity, service exposure, persistent storage, custom resources, and operators • Working knowledge of VXLAN EVPN fabrics • Knowledge of network segmentation and firewall policy design, including zone-based or microsegmentation approaches • Familiarity with the NetOps/NetSecOps operating split in data-center security • Experience driving diagnosis and remediation through a customer’s own team in environments with no direct access • Experience with NX-OS automation and APIs such as NX-API, NETCONF/RESTCONF, gNMI, or Ansible • Ability to communicate incident status, root cause, and remediation clearly to technical and executive audiences, verbally and in writing • CCNP Data Center, CCNP Enterprise, DevNet Professional, CCIE Data Center, CCIE Enterprise, or DevNet Expert is a plus, not required

🏖️ Benefícios

• Medical, dental and vision insurance • 401(k) plan with Cisco matching contribution • Paid parental leave • Short- and long-term disability coverage • Basic life insurance • Restricted stock unit grants may be available, subject to continued employment and vesting periods • 10 paid holidays per full calendar year • 1 floating holiday for non-exempt employees • 1 paid day off for employee’s birthday • Paid year-end holiday shutdown • 4 paid days off for personal wellness • 16 days of paid vacation time per full calendar year for non-exempt employees • Flexible vacation time off program with no defined limit for eligible exempt employees • 80 hours of sick time off provided on hire and each January 1st thereafter • Up to 80 hours of unused sick time carried forward • Additional paid time away for critical or emergency family issues • Optional 10 paid volunteer days per full calendar year • Annual bonuses may be available for non-sales roles • Opportunities to grow and build on a global scale

Candidatar-se

Vagas Similares

🕒 6 dias atrás

Ad Hoc LLC

501 - 1000

💼 Consultoria

🏥 Saúde

📦 Logística

DevOps Engineer III building AWS infrastructure, Terraform modules, and GitLab CI/CD platforms for government digital services. Supporting secure container deployments, observability, incident response, and federal compliance.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $125.000 - $142.000 / ano

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 6 dias atrás

The Access Group

5001 - 10000

💼 Consultoria

🏥 Saúde

🏨 Hospitalidade

Senior Site Reliability Engineer operating secure, scalable Azure platforms for Access, a business management software provider. Leading incident remediation, Kubernetes, Terraform, observability, compliance, and DevOps improvements.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $165.000 - $185.000 / ano

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 6 dias atrás

AAA Life Insurance Company

501 - 1000

⚕️ Seguro de Saúde

💸 Finanças

🛡️ Seguros

Senior DevOps Engineer modernizing AAA Life’s cloud-native infrastructure and middleware. Designing CI/CD, automation, observability, security, and disaster recovery for life insurance systems.

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 6 dias atrás

DraftKings Inc.

1001 - 5000

🎮 Jogos

⚽ Esportes

👥 B2C

Database Reliability Engineer strengthening DraftKings’ database infrastructure for real-time sports betting and gaming. Automating Kubernetes platforms, improving performance, and building self-healing systems.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $112.000 - $140.000 / ano

⏰ Tempo Integral

🟢 Júnior

🟡 Pleno

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🚫👨‍🎓 Sem graduação necessária

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 6 dias atrás

Net Health

501 - 1000

🏥 Saúde

☁️ SaaS

🤖 Inteligência Artificial

DevOps Engineer designing secure AWS platforms and CI/CD automation for Net Health’s healthcare SaaS products. Owning cloud architecture, database performance, observability, security, and cost optimization.

🗣️🇺🇸🇬🇧 Inglês obrigatório