Network Automation, Reliability Engineer

Job not on LinkedIn

🔥 5 hours ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Ontrac Solutions

Ontrac Solutions

11 - 50 employees

Founded 2010

🤖 Artificial Intelligence

💼 Consulting

🤝 B2B

Artificial Intelligence • Consulting • B2B

Ontrac Solutions is an AI-first technology services firm that helps organizations design, implement, and scale AI-driven systems, cloud architectures, and enterprise platform integrations. They provide AI strategy and generative AI implementation (including RAG and intelligent search), cloud migration and optimization (multi-cloud architecture, Kubernetes, Terraform, GitOps, and FinOps), data and integration engineering, CRM/CMS/e-commerce platform development, and embedded technical talent / staff augmentation to move AI from experimentation into production. With innovation hubs in Chicago and Karachi, Ontrac focuses on delivering production-ready AI and cloud solutions that drive measurable business outcomes.

📋 Description

• Design, write, and maintain Python tooling that provisions, validates, and audits network devices across the global fleet • Build model-driven provisioning libraries using Jinja2 templating with NETCONF/YANG or REST APIs to codify golden configurations and eliminate configuration drift • Develop automated remediation for recurring failure classes triggered by syslog and telemetry • Identify recurring manual runbook steps and replace them with tested, reviewed code • Operate and scale data center fabrics, backbone links, and out-of-band management networks across data centers and POP sites • Design and tune BGP policy to control path selection and eliminate single-carrier points of failure • Support capacity expansion and site turn-ups, including high-level design, port channelization, optics and cabling standards, and hardware qualification • Execute zero-downtime production change management, including hitless migrations and staged rollbacks • Operationalize multi-vendor streaming telemetry and tune subscription jobs • Build and maintain observability for hop-by-hop path tracing, multi-layer fault isolation, and root cause analysis • Contribute to a real-time network source of truth aggregating BGP, link-state, and drain-state data • Participate in a 24x7 on-call rotation, lead incident response, write RCAs, and drive follow-up automation • Maintain and optimize firewall and ACL policy across multi-vendor platforms • Support SIRT/PSIRT CVE remediation through automated regression testing and config-as-code pipelines • Author and maintain technical documentation, including designs, runbooks, and API contracts

🎯 Requirements

• Demonstrated, sustained Python development in a production network or infrastructure environment • Production tooling experience for config generation and validation, API integrations, telemetry collectors, automated remediation, or test harnesses • Comfort with modules, packaging, testing, code review, and version control • Production BGP experience, including policy, path selection, and multihoming • Experience with IS-IS or OSPF, ECMP, and VXLAN/EVPN or MPLS overlays • Hands-on operations across at least two of Arista EOS, Juniper Junos (QFX/SRX/PTX/MX), or Cisco IOS-XR/NX-OS • Experience with Ansible and Jinja2 • Experience with NETCONF/YANG, RESTCONF, or vendor REST APIs • Experience with gNMI/gRPC streaming telemetry, OpenConfig models, SNMP, flow telemetry, and dashboarding/alerting • Linux experience, including networking stack, packet capture, systemd services, and shell scripting • Git-based workflows with peer review and CI pipelines such as Jenkins, GitLab CI, or GitHub Actions • Experience holding a 24x7 production on-call rotation, including incident command and root cause analysis • Roughly 2–5 years of experience in network production, network reliability, or network automation • Must be located in the United States and authorized to work in the US • Preferred: out-of-band network experience, data center or POP build-out, MACsec/IPsec, 802.1X/NAC, optical or transport technologies, NetBox or in-house source of truth, LLM or agentic on-call tooling, and a master's degree

Apply Now

Similar Jobs

🕒 Yesterday

MKS2 Technologies

201 - 500

💼 Consulting

🏥 Healthcare

📦 Logistics

Mid-level DevSecOps practitioner securing AWS cloud platforms for government clients. Supporting Kubernetes, CI/CD, GitOps, automation, and cloud modernization remotely.

🕒 6 days ago

Outpost

51 - 200

🚗 Transport

📦 Logistics

🏠 Real Estate

Site Reliability Engineer ensuring system reliability for a logistics platform. Working on backend/API, monitoring, and proactive incident response with a high-performance team.

🕒 July 17

Leland

11 - 50

💼 Consulting

📣 Marketing

📚 Education

Site Reliability & Engineering Coach for a remote platform connecting people to career experts. Focusing on practical skill development in site reliability engineering and career support.

🕒 June 25

Black Pearl Consult

11 - 50

💼 Consulting

🎯 Recruiter

Junior DevOps Engineer supporting CI/CD pipelines and cloud infrastructure at Black Financial Consult. Ideal for early-career professionals interested in automation and modern software delivery practices.

🕒 June 24

Black Pearl Consult

11 - 50

💼 Consulting

🎯 Recruiter

DevOps Engineer focusing on automating deployment pipelines and managing cloud infrastructure for technology company. Collaborating with engineering and cloud teams to improve deployment speed and reliability.