Staff Network Reliability Engineer – RAN Operations

🕒 2 dias atrás

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $150.000 - $162.000 / ano

⏰ Tempo Integral

🔴 Especialista

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

Candidatar-se
Encontrar Vagas Remotas Similares

📊 Verifique sua pontuação de currículo para esta vaga

Melhore suas chances de conseguir uma entrevista verificando sua pontuação de currículo antes de se candidatar.

Logo of Skylo

Skylo

51 - 200 funcionários

Fundada em 2017

📡 Telecomunicações

🤝 B2B

💰 $30.000.000 Venture Round - Skylo em 2025-02

Telecommunications • B2B

A Skylo é uma empresa que oferece serviços de conectividade global, focando na conexão de dispositivos IoT e edge por meio de redes celulares habilitadas por satélite. O site da empresa destaca mapas de cobertura, dispositivos certificados, um programa de certificação, documentação para desenvolvedores, white papers, estudos de caso e um ecossistema de parceiros — indicando uma oferta B2B voltada para desenvolvedores e parceiros, que permite que dispositivos permaneçam conectados onde a cobertura celular tradicional é limitada. A Skylo se posiciona como um provedor de conectividade em telecomunicações para empresas, fabricantes de dispositivos e parceiros que desenvolvem soluções que precisam de ampla cobertura geográfica.

Descrição

• Own 24×7 NTN RAN health across Skylo's production network: gNodeB/eNodeB status, node availability, attach and access success rates, RRC setup and abnormal release behavior, RF integrity (SINR, RSSI, noise floor stability), and RAN-layer SLA compliance. • Monitor and triage RAN alarms using OSS dashboards, Grafana/inhouse telemetry, and NTN-specific alarm patterns — distinguish transient RF fluctuations from systemic degradation before escalating or acting. • Execute and own RAN-domain runbooks for all fault categories: cell outage recovery, timing resynchronization, parameter rollback, restart procedures, and interference mitigation actions — without requiring engineering team involvement for covered fault classes. • Own timing and synchronization health: GNSS lock status, Octoclock/reference clock drift, PTP/SyncE alignment, and NTN-specific timing compensation — detect and resolve timing anomalies before they cascade into access failures. • Monitor beam-level and cell-level performance across the NTN coverage area; identify underperforming beams and initiate optimization or escalation based on defined criteria. • Serve as the L3 escalation authority for all RAN-domain incidents: take ownership from the Incident Manager, diagnose at the RF and protocol level using eNodeB/gNodeB logs, RF telemetry, PRACH/MSG1–MSG5 analysis, RRC traces, and KPI correlation, and deliver a resolution or a decision-grade root cause. • Lead RAN-domain troubleshooting bridges: command the technical investigation, direct vendor and engineering participants, correlate signals across RU, DU, CU-CP, CU-UP, and timing subsystems, and drive the bridge to a documented resolution or a clear engineering handoff. • Diagnose and resolve RAN failure modes: SINR anomalies, beam coverage gaps, preamble failures (MSG1–MSG5), RRC setup failures, abnormal RRC releases, CW/narrowband/adjacent-carrier interference, cell unavailability, eCPRI link failures, and DU/CU software faults. • Engage RAN vendors with technical specificity: reproduce failures with log evidence and RF traces, own the vendor ticket lifecycle, enforce SLA response commitments, and escalate vendor delays with full impact context. • Participate in the global 24×7 on-call rotation as the RAN domain escalation tier — reachable within defined SLA windows for Sev 1 events; function as the technical decision-maker, not the first responder. • Own RAN-domain RCA end-to-end: lead the post-incident investigation, document the complete causal chain from triggering RF condition or hardware event through downstream subscriber impact, and deliver systemic action items with owners, timelines, and measurable success criteria. • Deliver Initial RCA documentation within defined SLA windows post-incident closure; own the final RCA through engineering review and sign-off. • Identify systemic RAN failure patterns — recurring interference sources, timing drift trends, vendor software regressions, hardware batch faults — and translate them into engineering requirements with clear impact, scope, and acceptance criteria. • Contribute to the weekly and monthly Network Performance Report: RAN availability by beam/cell, attach success rates, MTTR by fault category, top recurring issues, and SLA deviation analysis.

🎯 Requisitos

• 8–10+ years of experience in RAN engineering and operations in a production 24×7 environment — carrier or vendor side, with direct ownership of live LTE/5G RAN infrastructure. NTN or satellite RAN experience strongly preferred. • Deep 3GPP RAN expertise: TS 38.300, NR/LTE-NTN procedures, NB-IoT/LTE-M/5G-NR air interfaces, UE random-access procedures (PRACH/MSG1–MSG5), RRC state machine, and beam management. • Extensive understanding of vRAN/ORAN architecture: 7-2x split functions, RU/DU/CU-CP/CU-UP component roles, eCPRI/CPRI interfaces, and virtualized RAN software systems. • Production RAN troubleshooting: demonstrated ability to diagnose SINR anomalies, beam coverage gaps, preamble failures, timing drift, interference conditions (CW, narrowband, adjacent carrier), and RU/DU/CU hardware and software faults using RF telemetry and gNB logs. • RF and timing domain knowledge: GNSS, PTP, SyncE, Octoclock/reference clock behavior, and NTN-specific timing compensation mechanisms. • Production observability: Prometheus/Grafana/in house tools for RAN KPI dashboarding; OSS alarm integration (SNMP, Pub/Sub, or equivalent). • Packet capture and trace analysis: proficiency with Wireshark or equivalent for L2/L3 call flows, PCAP analysis, and gNB trace decoding. • Kubernetes operational literacy: pod health monitoring for DU/CU workloads, kubectl proficiency, log aggregation and correlation. • Runbook authorship: ability to write RAN diagnostic procedures at the level where a less-experienced engineer can execute them independently. • Strong written and verbal communication: capable of delivering RCA documents, engineering escalations, and MNO-facing technical summaries.

🏖️ Benefícios

• Competitive compensation packages including a stock option-based equity program • Comprehensive benefits including medical, dental, vision, and retirement plan • Monthly allowances for wellness and education reimbursement • A generous time-off policy, holidays, and the opportunity to temporarily work abroad • A once-in-a-career opportunity to operate the world's first commercial, live direct-to-device satellite network • Access to a world-class team across software, hardware, chipsets, telecom, satellite, and network virtualization • Open, transparent, inclusive culture that blends Silicon Valley, Nordic, and South Asia characteristics

Candidatar-se

Vagas Similares

🕒 2 dias atrás

CDW

10.000+ funcionários

💼 Consultoria

🏥 Saúde

📚 Educação

Principal Consulting Engineer optimizing private cloud infrastructure built on VCF 9.0 for CDW. Driving standardization and automation across compute, storage, and networking management layers.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $164.000 - $240.800 / ano

💰 Post-IPO Equity em 2015-07

⏰ Tempo Integral

🔴 Especialista

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🦅 Patrocina Visto H1B

info

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 2 dias atrás

RTX

10.000+ funcionários

🚀 Aeroespacial

🎖️ Defesa

🏭 Manufatura

Site Reliability Engineer automating and improving reliability at Collins Aerospace. Collaborating with teams to tackle complex technical problems in flight operations.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $107.500 - $204.500 / ano

💰 $200.000 Grant - RTX em 2024-11

⏰ Tempo Integral

🔴 Especialista

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 2 dias atrás

Ad Hoc LLC

501 - 1000

💼 Consultoria

🏥 Saúde

📦 Logística

Staff DevOps Engineer at Ad Hoc shaping the long-term technical strategy and mentoring team members. Leading critical projects and ensuring compliance in delivering software for Veterans Affairs.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $130.000 - $150.000 / ano

⏰ Tempo Integral

🔴 Especialista

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 4 dias atrás

Cisco

10.000+ funcionários

🔧 Hardware

🔐 Segurança

🏢 Corporativo

Site Reliability Engineer focusing on building and maintaining cloud infrastructure for Cisco Meraki. Analyzing reliability, troubleshooting, and implementing solutions in a secure environment.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $167.700 - $245.200 / ano

⏰ Tempo Integral

🟠 Sênior

🔴 Especialista

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🦅 Patrocina Visto H1B

info

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 4 dias atrás

FluidStack

11 - 50

🤖 Inteligência Artificial

Principal Operations Engineer responsible for fleet reliability at Fluidstack's AI compute infrastructure. Defining availability targets and implementing corrective actions to improve operations.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $220.000 - $260.000 / ano

⏰ Tempo Integral

🔴 Especialista

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório