Senior Engineer, Cloud Infrastructure and Networking

🔥 12 hours ago

🇺🇸 United States – Remote

💵 $125k - $135k / year

⏰ Full Time

🟠 Senior

☁️ Cloud Engineer

👻 Ghost score 9%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Skylo

Skylo

51 - 200 employees

Founded 2017

📡 Telecommunications

🤝 B2B

💰 $30M Venture Round - Skylo on 2025-02

Telecommunications • B2B

Skylo is a company that provides global connectivity services focused on connecting IoT and edge devices through satellite-enabled cellular networks. The company's website highlights coverage maps, certified devices, a certification program, developer documentation, white papers, case studies, and a partner ecosystem — indicating a developer- and partner-focused B2B offering that enables devices to stay connected where traditional cellular coverage is limited. Skylo positions itself as a telecommunications connectivity provider for enterprises, device manufacturers, and partners building solutions that need broad geographic coverage.

📋 Description

• Own 24x7 cloud infrastructure health across Skylo’s hybrid production environment • Operate GCP public cloud and on-premise private cloud infrastructure • Monitor and triage infrastructure alarms using OSS dashboards, Grafana/VictoriaMetrics, GCP Cloud Monitoring, and Loki • Execute runbooks for GKE node recovery, pod eviction/rescheduling, PVC repair, database failover, Prometheus WAL recovery, ArgoCD drift remediation, and certificate rotation • Own the observability pipeline, including Prometheus, VictoriaMetrics, Grafana, OpenTelemetry, and alert routing • Maintain PostgreSQL replication, backups, restores, failover testing, query performance, and Redis operations • Ensure log aggregation pipeline health using Loki or ELK • Partner with Network Implementation on infrastructure changes and operational readiness • Serve as L3 escalation authority for Cloud Infrastructure incidents • Lead troubleshooting bridges and diagnose Kubernetes, storage, network, database, and GitOps failures • Participate in the global 24x7 on-call rotation • Define and maintain SLOs, track error budgets, reduce toil, and lead capacity planning • Own infrastructure root-cause analyses and post-incident action items • Author and maintain Cloud Infrastructure runbooks and SOPs • Validate operational readiness for infrastructure expansions, upgrades, and hardware deployments • Represent Cloud Infrastructure in architecture reviews and collaborate with NRE, security, platform engineering, and automation teams • Mentor Senior NREs in Kubernetes, storage, database reliability, observability, and escalation practices • Oversee ArgoCD, Helm, Terraform, and Ansible changes affecting production

🎯 Requirements

• 5+ years of infrastructure engineering, Site Reliability Engineering, or cloud operations in a production 24x7 environment • Direct on-call ownership for Kubernetes-at-scale environments • Deep Kubernetes expertise, including multi-cluster operations, node pool management, RBAC, network policies, PVCs, CSI drivers, CRD/operator patterns, and production cluster upgrades • Hands-on experience operating public cloud and on-premise/private cloud infrastructure • Production observability stack ownership using Prometheus, Grafana, VictoriaMetrics, OpenTelemetry, and Pub/Sub or equivalent • PostgreSQL streaming replication, backup/restore, failover procedures, and performance tuning • Redis cluster operations and persistence management • Production GitOps experience with ArgoCD or Flux CD • Helm chart authorship and version management • Terraform or Ansible for infrastructure provisioning • SRE fundamentals including SLO/SLI/SLA definition, error budget management, toil measurement, capacity planning, and on-call rotation design • Container runtime debugging, kernel-level performance analysis, storage subsystem troubleshooting, and network packet flow understanding • Ability to author runbooks executable independently by less-experienced engineers under incident pressure • Strong written and verbal communication for RCA documents, structured engineering escalations, and MNO-facing infrastructure summaries • Preferred: telecom or NTN workload experience • Preferred: Ceph, Rook, or equivalent distributed storage expertise • Preferred: KubeVirt, Harvester, or OpenStack experience • Preferred: BGP, VXLAN, EVPN, software-defined networking, and hardware load balancers • Preferred: Go or Python development • Preferred: FinOps experience • Preferred certifications: CKA, CKS, AWS Solutions Architect Professional, or Red Hat Certified Architect • Legally authorized to work in the U.S.

🏖️ Benefits

• Stock option-based equity program • Medical, dental, and vision benefits • Retirement plan • Monthly wellness allowance • Monthly education reimbursement • Generous time-off policy • Paid holidays • Opportunity to temporarily work abroad • Access to a world-class team across software, hardware, chipsets, telecom, satellite, and network virtualization • Flexible approach to work • Inclusive and diverse workplace culture

Apply Now

Similar Jobs

🔥 19 hours ago

OpenText

10,000+ employees

🏢 Enterprise

☁️ SaaS

🤝 B2B

Senior Cloud Services Project Manager leading client eDiscovery engagements for OpenText, a global information-management company. Managing legal-tech workflows, client delivery, and project billing.

🔥 20 hours ago

MariaDB

201 - 500

🏢 Enterprise

Senior Software Engineer building multi-cloud infrastructure and AI-enabled database applications for MariaDB Cloud. Owning production features, APIs, orchestration, and scalable relational database services.

🕒 Yesterday

Koniag Government Services

1001 - 5000

🏛️ Government

🎖️ Defense

💼 Consulting

Cloud Engineer supporting AWS WordPress infrastructure, Kubernetes CI/CD, and security for U.S. federal government agencies. Maintaining cloud environments, automation, monitoring, and compliance controls.

🕒 Yesterday

Hewlett Packard Enterprise

10,000+ employees

🏢 Enterprise

🔧 Hardware

🤖 Artificial Intelligence

Regional Sales Leader managing HPE sales teams and strategic accounts. Driving cloud and service-provider revenue, pipeline growth, executive partnerships, and talent development.

🕒 Yesterday

Hewlett Packard Enterprise

10,000+ employees

🏢 Enterprise

🔧 Hardware

🤖 Artificial Intelligence

Regional sales leader managing teams, pipelines, and strategic accounts for Hewlett Packard Enterprise’s edge-to-cloud technology business. Driving profitable growth, executive relationships, and sales performance.