Senior Software Engineer – Grafana Databases, Managed Services

🕒 May 7

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Grafana Labs

Grafana Labs

501 - 1000 employees

Founded 2014

🏢 Enterprise

☁️ SaaS

🤖 Artificial Intelligence

Enterprise • SaaS • Artificial Intelligence

Grafana Labs is a company that specializes in open-source observability technologies and solutions. It offers a comprehensive suite of tools for logging, metrics, tracing, and profile management with products like Grafana, Loki, Tempo, and Mimir. Their offerings are designed to help businesses visualize, monitor, and alert on data from various sources, providing capabilities such as anomaly detection, root cause analysis, and service level objective management using AI/ML insights. Grafana Labs provides both cloud-based and self-managed solutions, ideal for infrastructure, application, and frontend observability. Additionally, their platform supports integration with various data sources like Prometheus and OpenTelemetry, making them a key player in the observability and infrastructure monitoring space.

📋 Description

• Operating and evolving 100+ multi-cloud streaming clusters and related database infrastructure • Diagnosing and eliminating cross-layer failure modes (e.g., object storage latency, noisy neighbors, control-plane bottlenecks, query performance regressions, etc.) • Designing safe upgrade and rollout strategies at scale • Improving observability, automation, and operational ergonomics • Partnering closely with database and platform teams to ensure safe scaling, partitioning, consumer fan-out, and query performance • Working directly with distributed systems behavior, Kubernetes scheduling dynamics, storage engines, compression trade-offs, etc. • Serving as a primary escalation point and on-call for relevant incidents • Owning the relationship with all system vendors, including WarpStream Labs and others.

🎯 Requirements

• 6+ years of engineering experience, including meaningful time in SRE, platform engineering, production engineering, infrastructure engineering, or distributed systems roles. • Experience operating distributed systems in production (e.g., streaming systems, analytical databases, large-scale storage backends). Examples of these include Kafka, Redpanda, WarpStream, Postgres, ClickHouse, Snowflake, or Cassandra. • Strong Kubernetes experience in AWS, GCP, or Azure, and familiarity with infrastructure-as-code tooling (Helm, Terraform, Jsonnet, etc.). • Solid understanding of distributed systems design and large-scale system trade-offs. • Proficiency in at least one programming language (Go preferred, but not required). • Working knowledge of Linux internals, networking, cloud storage, and performance/scaling behavior. • Experience participating in blameless incident response and writing high-quality post-incident reviews. • Clear communicator who can collaborate across teams and work autonomously. • Curious, pragmatic, action-oriented, and kind (this is important!).

🏖️ Benefits

• 100% Remote, Global Culture • Scaling Organization – Tackle meaningful work in a high-growth, ever-evolving environment. • Transparent Communication – Expect open decision-making and regular company-wide updates. • Innovation-Driven – Autonomy and support to ship great work and try new things. • Open Source Roots – Built on community-driven values that shape how we work. • Empowered Teams – High trust, low ego culture that values outcomes over optics. • Career Growth Pathways – Defined opportunities to grow and develop your career. • Approachable Leadership – Transparent execs who are involved, visible, and human. • Passionate People – Join a team of smart, supportive folks who care deeply about what they do. • In-Person onboarding - We want you to thrive from day 1 with your fellow new ‘Grafanistas’ to learn all about what we do and how we do it. • Balance is Key - We operate a global annual leave policy of 30 days per annum. 3 days of your annual leave entitlement are reserved for Grafana Shutdown Days to allow the team to really disconnect.

Apply Now

Similar Jobs

🕒 May 6

Airbnb

5001 - 10000

👥 B2C

🛍️ eCommerce

Senior Software Engineer at Airbnb developing experiences for guests and hosts in EMEA. Leading technical projects and collaborating with global teams to optimize application functionalities.

Java

Kotlin

🕒 May 6

Vesta Software Group

201 - 500

🏢 Enterprise

🤝 B2B

Technical Lead developing reusable AI capabilities for Vesta Software Group across various business units. Leading hands-on engineering efforts and technical discovery while ensuring quality delivery.

AWS

Azure

Cloud

SDLC

🕒 May 6

The Audio Programmer

1 - 10

🤝 B2B

📚 Education

📱 Media

Full Stack Software Engineer focusing on building AI-powered features for music technology. Join a music technology company’s R&D team, collaborating with fellow engineers.

JavaScript

Node.js

React

TypeScript

🕒 May 5

Yelp

1001 - 5000

Backend engineer developing and supporting platforms for Yelp's Search systems. Collaborating with multiple teams on distributed systems and optimizing search-related infrastructure.

AWS

Distributed Systems

GraphQL

Java

NoSQL

Python

SQL

Unix

🕒 May 5

Jonas Software

1001 - 5000

🏢 Enterprise

☁️ SaaS

🤝 B2B

Technical Lead leading AI capability development for Vesta Software Group’s business units. Building production-grade AI solutions with a focus on collaboration and engineering excellence.

AWS

Azure

Cloud

SDLC