Senior Site Reliability Engineer – Network Observability

Job not on LinkedIn

🔥 0 minutes ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of BPCS, Comprehensive marketing solutions, ltd.

BPCS, Comprehensive marketing solutions, ltd.

1 - 10 employees

💼 Consulting

📣 Marketing

Consulting • Marketing

BPCS, Comprehensive marketing solutions, ltd. is a marketing agency founded in 2015 by four friends and experienced advertising professionals. The company offers integrated marketing, business consulting, advertising, and marketing strategies. Positioned as a 'big little agency,' BPCS combines the resources of larger firms with the personal touch and commitment of smaller agencies, providing responsive and dedicated services to meet complex marketing challenges.

📋 Description

• Support the operation and reliability of enterprise network observability platforms that ingest and analyze network telemetry • Administer and operate Linux and Windows virtual machines hosted in a cloud environment • Maintain compliance with security, configuration, and operational standards • Operate and support network observability platforms, primarily Syslog-NG and Trapd • Execute operating system, application, and security upgrades and patching • Investigate automated alerts, customer-reported incidents, and platform performance issues • Troubleshoot network observability configurations, software applications, and operating systems • Create and maintain processing rules using regular expressions • Support enterprise network performance and application monitoring platforms • Analyze traffic patterns, network telemetry, resource utilization, system performance, and security events with network and security engineers • Conduct system capacity planning and recommend scalability and reliability improvements • Automate recurring operational tasks using Bash, PowerShell, or Python • Deploy, configure, and manage cloud services for reliability, scalability, security, and cost efficiency • Implement DevOps practices including CI/CD pipelines, source control, and infrastructure as code • Perform business continuity and disaster recovery failover testing • Manage assigned projects and program components according to objectives and timelines • Participate in daily stand-ups and collaborate with engineering teams • Participate in an on-call rotation and provide incident response • Deliver compliance, incident-resolution, and project-delivery outcomes

🎯 Requirements

• Bachelor’s degree in computer science, computer engineering, information technology, or a related technical field, or equivalent professional experience • Five to seven years of enterprise experience in systems engineering, network engineering, site reliability engineering, or a related IT infrastructure role • At least five years of Linux system administration experience • At least three years of network engineering experience • At least three years of hands-on Syslog-NG experience • Strong understanding of SNMP, SNMP Traps, NetFlow, and gNMI • Strong knowledge of enterprise networking, including routing and switching protocols • Hands-on experience with Azure or a comparable cloud platform • Experience operating and troubleshooting network monitoring systems in a large enterprise environment • Experience with system capacity planning, configuration management, compliance audits, and performance analysis • Ability to investigate complex incidents involving infrastructure, operating systems, applications, and network telemetry • Strong project management, collaboration, and communication skills • Preferred: proficiency in Bash, PowerShell, or Python for automation • Preferred: strong proficiency with regular expressions • Preferred: experience with IBM SevOne Network Performance Manager, Broadcom AppNeta, or comparable observability platforms • Preferred: experience using source-control platforms and development workflows • Preferred: working knowledge of Ansible and Ansible playbooks • Preferred: intermediate knowledge of KQL, T-SQL, or comparable data-retrieval languages • Preferred: experience implementing CI/CD pipelines and infrastructure-as-code practices • Preferred: exposure to commercially available artificial intelligence platforms • Preferred: experience connecting network telemetry with AI-enabled workflows

🏖️ Benefits

• Medical, dental, and vision coverage • Flexible Spending Account (FSA) • 401(k) retirement plan • Competitive paid time off • Parental leave • Professional growth and development opportunities

Apply Now

Similar Jobs

🔥 0 minutes ago

phData

201 - 500

💼 Consulting

🏥 Healthcare

🏭 Manufacturing

Senior DevOps Engineer operating Snowflake and cloud AI platforms for phData, a data and AI consultancy. Leading managed-services reliability, incident response, automation, and client delivery.

🔥 39 minutes ago

Salve.Inno

11 - 50

💼 Consulting

📣 Marketing

📦 Logistics

Senior Site Reliability Engineer operating AWS and Kubernetes cloud platforms for mission-critical production services. Improving observability, automation, incident response, and reliability across engineering teams.

🔥 1 hour ago

NVIDIA

10,000+ employees

🏥 Healthcare

🏭 Manufacturing

🤖 Artificial Intelligence

Technical Marketing Engineer building NVIDIA enterprise infrastructure solutions for global marketing and sales teams. Designing reference architectures, demos, data center environments, and technical documentation.

🔥 2 hours ago

Astreya

1001 - 5000

💼 Consulting

📦 Logistics

📣 Marketing

Data center engineer designing POP infrastructure, rack layouts, power systems, and deployment documentation. Supporting Astreya’s global IT managed services through network infrastructure projects and vendor coordination.

🔥 4 hours ago

Planet Depos

201 - 500

⚖️ Legal

🤝 B2B

Lead DevOps Engineer owning compliant AWS environments, databases, CI/CD, and observability for legal-industry software. Mentoring platform engineers and enabling secure, reliable product delivery.