DevOps Engineer, Observability

Job not on LinkedIn

🔥 45 minutes ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Twilio

Twilio

5001 - 10000 employees

🔌 API

🤝 B2B

API • B2B

Twilio is a leading provider of cloud communications services that enables developers to build innovative communication solutions. Founded in 2008, Twilio has democratized access to communication channels such as voice, text, chat, video, and email through easy-to-use APIs. With headquarters in San Francisco and a global presence, Twilio empowers organizations of all sizes to engage effectively with their customers by integrating these communication capabilities into their applications.

📋 Description

• Lead the end-to-end architecture and delivery of key observability platform components, focusing on reliability, scalability, and usability • Drive consistency and quality across logs, metrics, traces, and continuous profiling • Serve as a technical advisor and mentor across the platform organization • Guide design decisions and align cross-team efforts with long-term architectural goals • Address high-cardinality telemetry, distributed tracing correlation, and compute cost insights • Collaborate with product teams, SREs, and developer experience groups • Design and build developer-friendly tooling and APIs for incident response, performance analysis, and platform debugging • Leverage open-source standards such as OpenTelemetry • Balance performance, cost, and user value across engineering teams

🎯 Requirements

• Proven expertise building and scaling observability systems • Experience with centralized S3-based data lakes, OpenTelemetry instrumentation, and ClickHouse-backed query engines • Proficiency in at least one modern programming language, such as Go, Python, or Java • Familiarity with high-cardinality data challenges and telemetry correlation techniques • Experience designing high-scale telemetry systems, such as Prometheus, ClickHouse, OpenTelemetry, or Kafka • Solid understanding of distributed systems and observability in complex microservice environments • Experience with AWS, Kubernetes, and infrastructure-as-code tools • Ability to provide architectural guidance and technical thought leadership across teams • Ability to make forward-looking technical decisions and lead through ambiguity • ClickHouse, Grafana Mimir, Athena, or equivalent systems experience is desired • Open-source observability contributions are desired • Cloud compute and telemetry FinOps tooling experience is desired

🏖️ Benefits

• Competitive pay • Generous time off • Ample parental leave • Wellness leave • Healthcare • Retirement savings program • Support for employee volunteering and donation efforts • Remote-first work

Apply Now

Similar Jobs

🕒 Yesterday

Arista Networks

1001 - 5000

🏢 Enterprise

📡 Telecommunications

Senior Site Reliability Engineer building and operating scalable, secure production infrastructure for Arista Networks’ cloud networking platforms. Automating operations, improving observability, and ensuring reliable deployments from Ireland.

AWS

Azure

Cloud

Distributed Systems

Docker

Google Cloud Platform

Grafana

Kubernetes

Linux

Postgres

Prometheus

Python

Shell Scripting

Spinnaker

Terraform

Unix

Go

🕒 2 days ago

Arista Networks

1001 - 5000

🏢 Enterprise

📡 Telecommunications

Site Reliability Engineer at Arista Networks focusing on building and operating critical production systems for scalability and reliability. Engaging in a collaborative remote role from Ireland.

AWS

Azure

Cloud

Distributed Systems

Docker

Google Cloud Platform

Grafana

Kubernetes

Linux

Postgres

Prometheus

Python

Shell Scripting

Spinnaker

Terraform

Unix

Go

🕒 July 28

Astreya

1001 - 5000

💼 Consulting

📦 Logistics

📣 Marketing

IT Infrastructure Support Engineer optimizing critical physical security systems for a global IT provider. Engineering automation tools for a reliable and scalable infrastructure environment.

Ansible

Chef

Cloud

Grafana

IoT

Kubernetes

Linux

Prometheus

Puppet

Python

Terraform

Go

🕒 July 27

Sardine

51 - 200

🔒 Cybersecurity

📋 Compliance

💳 Fintech

DevOps Engineer at Sardine improving infrastructure and tooling for a remote-first financial crime platform. Collaborating to ensure reliable, scalable, and cost-efficient systems.

AWS

Cloud

Distributed Systems

Google Cloud Platform

Kubernetes

Prometheus

Python

Terraform

Go

🕒 July 27

Holafly

501 - 1000

✈️ Travel

📡 Telecommunications

👥 B2C

DevSecOps Engineer at Holafly, designing secure GCP foundations and automating with Terraform. Safeguarding connectivity for millions of travelers through infrastructure and security enhancements.

Ansible

Cloud

Docker

Google Cloud Platform

Kubernetes

Python

Terraform