Technical Staff Member – Observability, Reliability

Job not on LinkedIn

🔥 6 minutes ago

🇧🇷 Brazil – Remote

⏰ Full Time

🔴 Lead

👻 Ghost score 12%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of avra

avra

1 - 10 employees

💼 Consulting

💸 Finance

🤝 B2B

Consulting • Finance • B2B

Avra is a venture capital firm dedicated to empowering growth stage founders by providing the essential support needed to scale their companies after Series A funding. With a commitment to building a prosperous future together, Avra focuses on strategic investments that foster innovation and entrepreneurship.

📋 Description

• Evolve the observability stack for logs, metrics, traces, and alerting • Ensure every cloud and on-premise dataplane reports its active release, health, heartbeat, logs, metrics, and usage to the control plane • Bring telemetry into customer clusters using outbound-only agent connections • Detect drift between desired state and actual state in each environment • Monitor deployment and runtime agent health • Provide visibility into ephemeral workloads such as Ray clusters running batch inference • Define SLOs, lead incident response and postmortems, and reduce MTTR, including customer-coordinated fixes • Reduce telemetry cost by eliminating redundant data and improving signal • Achieve 99.9% serving availability, reduce incidents and MTTR, minimize state drift, and ensure all agents are active and reporting

🎯 Requirements

• Deep experience with OpenTelemetry and observability backends • Hands-on practice with SLOs, error budgets, actionable alerting, and incident management • Strong experience with Kubernetes and infrastructure as code (Terraform / Helm) • Experience operating software in environments you don't fully control • Production-quality code and reviews, and a willingness to operate what you build • Nice to have: Shipping software to customer-hosted Kubernetes (e.g., Helm, outbound-only connectivity) • Nice to have: GCP or GKE, AWS or EKS • Nice to have: ML multi-node/multi-cluster workloads in production • Nice to have: Financial services or regulated environments

Apply Now