Senior Cloud Monitoring and Observability Engineer

🔥 0 minutes ago

🇺🇸 United States – Remote

💵 $107.9k - $195.1k / year

⏰ Full Time

🟠 Senior

👷🏻‍♀️ Engineer

🦅 H1B Visa Sponsor

infoinfo

👻 Ghost score 0%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Leidos

Leidos

10,000+ employees

Founded 1969

🏥 Healthcare

💼 Consulting

📦 Logistics

Healthcare • Consulting • Logistics

Leidos is a leading systems integrator in science, technology, and engineering, providing solutions that transform and enable the missions of its customers. The company operates across various markets, including aviation, defense, energy, government, healthcare, intelligence, science, and space. Leidos is involved in AI, digital modernization, cyber operations, and integrated and mission software systems. With a commitment to diversity, equity, inclusion, and sustainability, Leidos also engages in charitable efforts and community enrichment programs. Additionally, it contributes to developing solutions for counter-unmanned aerial systems and electric vehicle infrastructure for military applications.

📋 Description

• Engineer, operate, maintain, and continuously improve enterprise monitoring and observability capabilities across hybrid infrastructure, cloud, and container platforms • Manage dashboards, monitors, metrics, logs, APM, synthetic monitoring, tagging, integrations, and related platform capabilities • Assess monitoring coverage, identify visibility gaps, and coordinate onboarding or remediation with technical teams • Maintain monitoring coverage across Windows, Linux, cloud, OpenShift/Kubernetes, virtualized, database, network, storage, middleware, and application environments • Support monitoring for Red Hat OpenShift, Kubernetes, OpenShift Virtualization, and OpenShift-hosted virtual machines • Configure and troubleshoot monitoring agents, integrations, collectors, APIs, and platform components • Build and maintain tagging, metadata, dashboards, alerts, service health views, and operational reporting • Automate monitoring deployment, configuration, tagging, onboarding, upgrades, and integrations using Ansible, APIs, scripting, CI/CD, or infrastructure-as-code • Integrate observability platforms with ServiceNow, notification systems, on-call workflows, and enterprise operational systems • Use telemetry to troubleshoot performance and availability issues, support incident response and root-cause analysis, and recommend remediation • Correlate infrastructure, application, platform, and dependency telemetry to identify degradation and recurring issues • Partner with Operations and engineering teams to improve coverage, alert quality, service visibility, incident detection, escalation, and response • Analyze telemetry and historical trends to identify capacity risks, recurring issues, monitoring gaps, and improvement opportunities • Develop performance, availability, capacity, and monitoring coverage reporting for technical and leadership stakeholders • Maintain monitoring standards, technical documentation, configuration guidance, and operational procedures

🎯 Requirements

• BS degree and 8-12 years of prior relevant experience, or Master's degree with 6-10 years of prior relevant experience; additional relevant experience may be considered in lieu of degree requirements where permitted by contract • Strong hands-on experience engineering and operating enterprise monitoring or observability platforms • Production experience monitoring Windows and Linux infrastructure and Kubernetes or Red Hat OpenShift environments • Experience deploying, configuring, upgrading, and troubleshooting monitoring agents, integrations, dashboards, alerts, tagging, and operational reporting • Experience automating monitoring deployment or administration using Ansible, APIs, scripting, CI/CD pipelines, infrastructure-as-code, or similar technologies • Experience integrating monitoring or observability platforms with ITSM systems such as ServiceNow • Strong troubleshooting and dependency-analysis skills across infrastructure, applications, networks, platforms, and services • Ability to analyze technical telemetry, identify monitoring or performance gaps, and translate findings into actionable recommendations • Ability to communicate technical findings and recommendations to technical teams, project leadership, and customer stakeholders • Must meet applicable contract citizenship and work authorization requirements and be able to obtain and maintain SEC Public Trust or other required clearance • Preferred: Datadog experience; ScienceLogic SL1, SolarWinds, Dynatrace, New Relic, Splunk Observability, LogicMonitor, Prometheus/Grafana, or comparable platforms • Preferred: application performance monitoring, distributed tracing, OpenTelemetry, Datadog APM, Log Management, Synthetic Monitoring, RUM, Network Performance Monitoring, Database Monitoring • Preferred: Red Hat OpenShift Virtualization, CNV, KubeVirt, Microsoft Azure or AWS, Terraform, monitoring-as-code, API-driven deployment, FISMA, FedRAMP, NIST, Datadog/AWS/Azure/Red Hat OpenShift/Terraform/ITIL certifications

Apply Now

Similar Jobs

🔥 18 minutes ago

AeroVect

11 - 50

📦 Logistics

💼 Consulting

🚀 Aerospace

Senior Software Safety Engineer advancing safety for AeroVect’s autonomous ground-handling platforms. Performing ISO 26262 analyses, deriving requirements, and supporting airport deployment.

🔥 28 minutes ago

Sargent & Lundy

1001 - 5000

🏗️ Construction

🎖️ Defense

⚡ Energy

Lead I&C engineer modernizing nuclear power plants with digital control technologies. Managing multidisciplinary projects, nuclear control-system design, vendors, and client execution.

🔥 3 hours ago

1mind

11 - 50

🤖 Artificial Intelligence

☁️ SaaS

🤝 B2B

Prompt and context engineer designing LLM subagents and agentic workflows. Building 1mind’s autonomous customer-experience platform for next-generation revenue teams.

🔥 3 hours ago

Instacart

1001 - 5000

🍽️ Food & Beverage

📦 Logistics

🛍️ eCommerce

Senior Detection Engineer building threat detections for Instacart’s grocery delivery platform. Automating response and investigating adversary activity across cloud, endpoint, container, and SaaS environments.

🔥 3 hours ago

Ulteig

1001 - 5000

💼 Consulting

📦 Logistics

🏗️ Construction

Transmission Line Engineer designing 69kV–500kV transmission lines for Ulteig, an engineering firm transforming North America’s critical infrastructure. Developing construction documents and coordinating multidisciplinary projects.