Associate Director, Observability and Service Reliability

🔥 12 hours ago

🇨🇦 Canada – Remote

⏰ Full Time

🟠 Senior

👔 Director

👻 Ghost score 10%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Kyndryl

Kyndryl

10,000+ employees

Founded 2021

💼 Consulting

📦 Logistics

🏥 Healthcare

Consulting • Logistics • Healthcare

Kyndryl is a leading IT infrastructure services provider, serving thousands of enterprise customers worldwide. The company specializes in designing, building, managing, and modernizing complex, mission-critical information systems. Kyndryl offers a range of services including IT consulting, cloud services, cybersecurity, data and AI solutions, and digital workplace transformation. With a strong focus on innovation, partnerships, and co-creation, Kyndryl helps businesses tackle IT complexity and drive operational excellence. The company operates across various industries such as automotive, healthcare, banking, and more, providing expertise and solutions to address industry-specific challenges. Kyndryl's global network and strategic alliances empower enterprises to adapt to the evolving technology landscape, ensuring their essential systems are reliable and efficient.

📋 Description

• Own and mature the enterprise observability and service reliability strategy • Define enterprise standards for monitoring applications, infrastructure, cloud platforms, networks, endpoints, APIs, databases, middleware, and critical technology services • Establish expectations for metrics, logs, traces, events, synthetic monitoring, real user monitoring, digital experience, service health, and business transaction visibility • Identify monitoring gaps, redundant capabilities, excessive alerting, and opportunities to improve visibility • Move the organization toward proactive, predictive, and automated operations • Establish and mature the service reliability framework • Partner with technical service owners to define monitoring requirements and service reliability measures • Provide enterprise technical authority and architectural guidance for observability, monitoring, and service reliability • Develop monitoring patterns, reference architectures, standards, and reusable capabilities • Evaluate emerging observability, AIOps, automation, analytics, and service reliability capabilities • Lead strategy, architecture, governance, adoption, and optimization of Dynatrace, Nexthink, and related monitoring platforms • Manage strategic technology and vendor relationships • Drive Dynatrace adoption across applications, infrastructure, cloud, and digital services • Lead Nexthink and Digital Employee Experience strategy, including proactive remediation and automation • Improve event management, signal quality, event correlation, anomaly detection, automated diagnostics, and proactive remediation • Integrate observability platforms with ITSM, incident management, automation, collaboration, configuration, and operational data platforms • Establish governance for monitoring standards, tooling, integrations, data quality, licensing, and adoption • Develop executive and operational reporting on service health, reliability, performance, and user experience • Lead and develop observability, monitoring, reliability, and platform engineering professionals • Build partnerships across application, platform, infrastructure, cloud, network, cybersecurity, Digital Workplace, DevOps, SRE, Service Management, and business teams • Partner with Incident, Problem, Change, Major Incident Management, and Operational Resilience teams to improve detection, recovery, and prevention

🎯 Requirements

• Bachelor’s degree in Information Technology, Computer Science, Engineering, or a related discipline, or equivalent professional experience • 8 or more years of experience in observability, application performance, service reliability, infrastructure, cloud, engineering, or enterprise technology operations • 3 or more years of technical leadership or people leadership experience • Experience designing and operating monitoring and observability capabilities in complex enterprise environments • Experience with Dynatrace and familiarity with Nexthink or comparable Digital Employee Experience technologies • Knowledge of observability technologies, monitoring architectures, application performance management, infrastructure monitoring, event management, logging, tracing, synthetic monitoring, and user experience monitoring • Understanding of modern enterprise applications, cloud technologies, containers, APIs, networks, databases, infrastructure, and distributed architectures • Experience defining service health, reliability measures, Service Level Indicators, Service Level Objectives, and operational performance standards • Ability to influence senior technical leaders and translate complex technical concepts into business risk and operational outcomes • Strong leadership, architecture, analytical, problem-solving, communication, and stakeholder management skills • Preferred: Advanced Dynatrace experience in a large enterprise environment • Preferred: Experience implementing or scaling Nexthink • Preferred: Experience with ServiceNow and enterprise event management platforms • Preferred: Experience with Site Reliability Engineering, DevOps, AIOps, automation, OpenTelemetry, cloud native monitoring, and modern observability architectures • Preferred: Experience establishing enterprise monitoring standards or observability reference architectures • Preferred: Experience within complex, global, or highly regulated enterprise environments

🏖️ Benefits

• Flexible, supportive environment • Well-being prioritized • Hybrid-friendly culture • Be Well programs supporting financial, mental, physical, and social health • Personalized development goals • Continuous feedback • Certifications with Microsoft, Google, and Amazon • Coaching and hands-on learning experiences • Cutting-edge learning opportunities • Career-path tools and access to in-demand skills

Apply Now

Similar Jobs

🕒 August 19

argenx

1001 - 5000

🏥 Healthcare

💼 Consulting

🧬 Biotechnology

Associate Director leading MyPATH patient support programs for argenx, a global immunology biotech developing autoimmune therapies. Driving patient access, vendor operations, analytics, and cross-functional execution across Canada.

🕒 August 18

The Nature Conservancy

5001 - 10000

💼 Consulting

✈️ Travel

🤲 Charity

Associate Director securing major gifts for Nature United’s conservation mission. Managing donor portfolios, fundraising strategies, and cross-border philanthropic relationships.

🇨🇦 Canada – Remote

💵 $104k - $112k / year

⏰ Full Time

🟠 Senior

👔 Director

🕒 August 11

Anthesis Group

1001 - 5000

🏗️ Construction

📦 Logistics

💼 Consulting

Associate Director leading complex life cycle assessments for Anthesis, a global sustainability consultancy. Advising clients on carbon footprints, circular systems, emerging technologies, and sustainable product decisions.

🇨🇦 Canada – Remote

💵 $125k - $148k / year

💰 Private Equity Round on 2023-09

⏰ Full Time

🟠 Senior

👔 Director

🕒 August 4

Everest Clinical Research

501 - 1000

🏥 Healthcare

💼 Consulting

📦 Logistics

Associate Director leading R- and Python-based open-source applications for Everest, a clinical research CRO. Governing compliant automation, CDISC standards, AI-assisted development, and SDLC practices.

🇨🇦 Canada – Remote

💵 $170k - $190k / year

⏰ Full Time

🟠 Senior

👔 Director

🕒 July 31

Perseus Group, Constellation Software

10,000+ employees

🤝 B2B

☁️ SaaS

Director, Corporate Development responsible for nurturing corporate relationships for software acquisitions. Position entails extensive networking and research to support acquisition strategies.

🇨🇦 Canada – Remote

💵 $75.6k - $92.4k / year

⏰ Full Time

🟠 Senior

👔 Director