Search Remote Jobs

Senior Network Reliability Engineer

🔥 0 minutes ago

🇺🇸 United States – Remote

💵 $135k - $190k / year

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🦅 H1B Visa Sponsor

info
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Group 1001

Group 1001

501 - 1000 employees

💼 Consulting

🏥 Healthcare

💸 Finance

Consulting • Healthcare • Finance

Group 1001 is a collective focused on empowering companies and communities through innovative financial solutions and strategic partnerships. The company provides simple and accessible insurance and annuity products designed to help individuals manage and grow their savings. Their online investing platform offers users a digital means to control their financial futures. Group 1001 also invests in partnerships that enhance community development through education and sports. With a tech-driven culture, they are transforming the insurance industry and helping communities thrive.

📋 Description

• Define network-platform SLOs and error budgets, using them to gate changes • Lead postmortems focused on permanent remediation • Move network state into declarative, version-controlled code using Terraform or Pulumi, Ansible, Python, and Infra CI/CD • Build network policy as intent and implement Policy as Code controls • Design, deploy, and manage network infrastructure through Infrastructure as Code • Operate and extend the multi-account AWS Landing Zone, including Cloud WAN, Transit Gateway, IPAM, private DNS, and telemetry pipelines • Build platform abstractions for correctly onboarding accounts and services • Engineer Kubernetes networking, service mesh, eBPF observability, and cloud-tier identity integrations • Build structured alerts, runbooks, and operator and end-user monitoring dashboards • Provide technical guidance to junior engineers and mentor the team • Maintain and configure routers, switches, and firewalls at data centers and offices when required • Serve as an escalation point for network incidents and troubleshoot through root-cause analysis • Author runbooks and SOPs and package routine work for L1/L2 handoff • Coordinate reliability practices across Data Platforms, NOC/SOC, and Cyber Security

🎯 Requirements

• Deep understanding of TCP/IP, BGP, OSPF, VPNs, and SD-WAN architecture • Proven production experience with Terraform state management and modules, Ansible playbooks and roles, or similar • Proficiency in Python for automation and API interaction, or similar • Hands-on experience with Cloudflare, Zscaler, and/or enterprise firewalls • Experience configuring monitoring tools such as Datadog, Prometheus, or Grafana for meaningful alerts and dashboards • Service mesh experience with Istio, Linkerd, Consul Connect, or Cilium (nice to have) • eBPF-based observability experience with Hubble or Pixie (nice to have) • AWS multi-account landing-zone tooling experience, such as AFT or Control Tower (nice to have) • Policy as Code experience with OPA/Rego, Sentinel, or Cilium NetworkPolicy (nice to have) • Strong belief in documentation-first practices • Mindset focused on automating repetitive tasks and reducing toil • Willingness to handle physical hardware tasks when required while maintaining a software-centric engineering mindset

🏖️ Benefits

• Comprehensive health, dental, and vision insurance plan options • Basic and Supplemental Life Insurance • Short- and Long-Term Disability • Employee Assistance Program • Wellness programs • 401K plan with matching contributions by the Company • Supportive work environment where employee differences are valued

Apply Now

Similar Jobs

🔥 3 hours ago

PathAI

501 - 1000

🏥 Healthcare

💼 Consulting

📦 Logistics

Senior/Staff SRE designing and operating secure on-premises and hybrid-cloud data centers for PathAI’s AI-powered pathology platform. Improving reliability, automation, observability, and incident response for machine-learning infrastructure.

🔥 4 hours ago

Ardent

51 - 200

💼 Consulting

🎖️ Defense

📦 Logistics

DevSecOps Engineer securing cloud products and services for Ardent’s federal national security and defense missions. Automating deployments, vulnerability mitigation, and enterprise system architecture.

🔥 4 hours ago

Hexion Inc.

1001 - 5000

🚘 Automotive

🏗️ Construction

🏭 Manufacturing

Reliability Engineer improving asset performance across Hexion’s North American manufacturing plants. Leading failure elimination, maintenance optimization, and cross-site reliability standardization.

🔥 7 hours ago

Vontier

5001 - 10000

🚘 Automotive

⚡ Energy

🔧 Hardware

Hardware Reliability Engineer leading reliability planning, validation, and failure analysis for Gilbarco Veeder-Root fueling equipment. Improving durability, serviceability, and field performance across complex electromechanical products.

🔥 7 hours ago

Hexion Inc.

1001 - 5000

🚘 Automotive

🏗️ Construction

🏭 Manufacturing

Reliability Engineer reducing downtime and improving asset performance across Hexion’s North American manufacturing plants. Leading RCA, maintenance strategy, CMMS execution, KPI analysis, and cross-site reliability standardization.