Customer Reliability Engineer

🕒 6 days ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Cisco

Cisco

10,000+ employees

Founded 1984

🔧 Hardware

🔐 Security

🏢 Enterprise

Hardware • Security • Enterprise

Cisco is a multinational technology company that provides networking hardware, software, and services to enterprises, service providers, and governments. It builds routers, switches, optical transceivers, programmable silicon, and edge computing platforms, and offers security, collaboration (Webex), observability, and AI-enabled software and support services to help organizations design, operate, and secure large-scale networks and data centers. Cisco also delivers professional services, training, and cloud-managed solutions to support digital transformation and AI-ready infrastructure.

📋 Description

• Own Cisco Hypershield cases escalated from Cisco TAC through resolution, engaging customers directly as needed • Diagnose complex production failures across the N9300 Smart Switch fabric and on-premises Kubernetes controller • Localize faults across switching and forwarding, security services and enforcement, and control-plane layers • Understand customer architectures and configurations and diagnose failures in unfamiliar production environments • Reproduce customer failures, partner with engineering to drive fixes, and own the fix back to the customer • Convert individual cases into systemic improvements such as runbooks, diagnostics, knowledge-base content, and product feedback • Build the team’s proactive view of customer health through monitoring, tooling, and reliability practices

🎯 Requirements

• Bachelor’s degree plus 8 years of experience, Master’s degree plus 6 years, or equivalent industry experience • Experience supporting enterprise customers in an escalation capacity • Experience diagnosing and resolving complex production incidents under SLA pressure in unfamiliar environments • Experience operating and troubleshooting Cisco Nexus / NX-OS in production; equivalent depth on another major vendor accepted • Experience localizing failures across layered data-center architectures spanning switching/forwarding, services/enforcement, and control-plane domains • Linux operations experience at the command line, including production troubleshooting • Working exposure to containers or Kubernetes • Experience with packet capture and flow-telemetry analysis, including NetFlow/IPFIX • Working knowledge of enterprise virtualization and troubleshooting VM-based appliance deployments; vSphere is the current deployment target • Operational Kubernetes and Helm proficiency, including TLS certificates, service-account authentication, API-server connectivity, service exposure, persistent storage, custom resources, and operators • Working knowledge of VXLAN EVPN fabrics • Knowledge of network segmentation and firewall policy design, including zone-based or microsegmentation approaches • Familiarity with the NetOps/NetSecOps operating split in data-center security • Experience driving diagnosis and remediation through a customer’s own team in environments with no direct access • Experience with NX-OS automation and APIs such as NX-API, NETCONF/RESTCONF, gNMI, or Ansible • Ability to communicate incident status, root cause, and remediation clearly to technical and executive audiences, verbally and in writing • CCNP Data Center, CCNP Enterprise, DevNet Professional, CCIE Data Center, CCIE Enterprise, or DevNet Expert is a plus, not required

🏖️ Benefits

• Medical, dental and vision insurance • 401(k) plan with Cisco matching contribution • Paid parental leave • Short- and long-term disability coverage • Basic life insurance • Restricted stock unit grants may be available, subject to continued employment and vesting periods • 10 paid holidays per full calendar year • 1 floating holiday for non-exempt employees • 1 paid day off for employee’s birthday • Paid year-end holiday shutdown • 4 paid days off for personal wellness • 16 days of paid vacation time per full calendar year for non-exempt employees • Flexible vacation time off program with no defined limit for eligible exempt employees • 80 hours of sick time off provided on hire and each January 1st thereafter • Up to 80 hours of unused sick time carried forward • Additional paid time away for critical or emergency family issues • Optional 10 paid volunteer days per full calendar year • Annual bonuses may be available for non-sales roles • Opportunities to grow and build on a global scale

Apply Now

Similar Jobs

🕒 6 days ago

Ad Hoc LLC

501 - 1000

💼 Consulting

🏥 Healthcare

📦 Logistics

DevOps Engineer III building AWS infrastructure, Terraform modules, and GitLab CI/CD platforms for government digital services. Supporting secure container deployments, observability, incident response, and federal compliance.

🇺🇸 United States – Remote

💵 $125k - $142k / year

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🕒 6 days ago

The Access Group

5001 - 10000

💼 Consulting

🏥 Healthcare

🏨 Hospitality

Senior Site Reliability Engineer operating secure, scalable Azure platforms for Access, a business management software provider. Leading incident remediation, Kubernetes, Terraform, observability, compliance, and DevOps improvements.

🇺🇸 United States – Remote

💵 $165k - $185k / year

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

AWS

Azure

Cloud

DNS

Firewalls

Java

Kubernetes

MS SQL Server

Puppet

Scala

SQL

Terraform

🕒 6 days ago

AAA Life Insurance Company

501 - 1000

⚕️ Healthcare Insurance

💸 Finance

🛡️ Insurance

Senior DevOps Engineer modernizing AAA Life’s cloud-native infrastructure and middleware. Designing CI/CD, automation, observability, security, and disaster recovery for life insurance systems.

🕒 6 days ago

DraftKings Inc.

1001 - 5000

🎮 Gaming

⚽ Sports

👥 B2C

Database Reliability Engineer strengthening DraftKings’ database infrastructure for real-time sports betting and gaming. Automating Kubernetes platforms, improving performance, and building self-healing systems.

🇺🇸 United States – Remote

💵 $112k - $140k / year

⏰ Full Time

🟢 Junior

🟡 Mid-level

⛑ DevOps & Site Reliability Engineer (SRE)

🚫👨‍🎓 No degree required

🕒 6 days ago

Net Health

501 - 1000

🏥 Healthcare

☁️ SaaS

🤖 Artificial Intelligence

DevOps Engineer designing secure AWS platforms and CI/CD automation for Net Health’s healthcare SaaS products. Owning cloud architecture, database performance, observability, security, and cost optimization.