
501 - 1000 employees
Founded 2014
🏢 Enterprise
☁️ SaaS
🤖 Artificial Intelligence
Enterprise • SaaS • Artificial Intelligence
Grafana Labs is a company that specializes in open-source observability technologies and solutions. It offers a comprehensive suite of tools for logging, metrics, tracing, and profile management with products like Grafana, Loki, Tempo, and Mimir. Their offerings are designed to help businesses visualize, monitor, and alert on data from various sources, providing capabilities such as anomaly detection, root cause analysis, and service level objective management using AI/ML insights. Grafana Labs provides both cloud-based and self-managed solutions, ideal for infrastructure, application, and frontend observability. Additionally, their platform supports integration with various data sources like Prometheus and OpenTelemetry, making them a key player in the observability and infrastructure monitoring space.
🔥 0 minutes ago
AWS
Azure
Cassandra
Cloud
Distributed Systems
Google Cloud Platform
Kafka
Kubernetes
Linux
Postgres
Terraform
Go
Improve your chances of getting an interview by checking your resume score before you apply.

501 - 1000 employees
Founded 2014
🏢 Enterprise
☁️ SaaS
🤖 Artificial Intelligence
Enterprise • SaaS • Artificial Intelligence
Grafana Labs is a company that specializes in open-source observability technologies and solutions. It offers a comprehensive suite of tools for logging, metrics, tracing, and profile management with products like Grafana, Loki, Tempo, and Mimir. Their offerings are designed to help businesses visualize, monitor, and alert on data from various sources, providing capabilities such as anomaly detection, root cause analysis, and service level objective management using AI/ML insights. Grafana Labs provides both cloud-based and self-managed solutions, ideal for infrastructure, application, and frontend observability. Additionally, their platform supports integration with various data sources like Prometheus and OpenTelemetry, making them a key player in the observability and infrastructure monitoring space.
• Operating and evolving 100+ multi-cloud streaming clusters and related database infrastructure • Diagnosing and eliminating cross-layer failure modes (e.g., object storage latency, noisy neighbors, control-plane bottlenecks, query performance regressions, etc.) • Designing safe upgrade and rollout strategies at scale • Improving observability, automation, and operational ergonomics • Partnering closely with database and platform teams to ensure safe scaling, partitioning, consumer fan-out, and query performance • Working directly with distributed systems behavior, Kubernetes scheduling dynamics, storage engines, compression trade-offs, etc. • Serving as a primary escalation point and on-call for relevant incidents • Owning the relationship with all system vendors, including WarpStream Labs and others.
• 6+ years of engineering experience, including meaningful time in SRE, platform engineering, production engineering, infrastructure engineering, or distributed systems roles. • Experience operating distributed systems in production (e.g., streaming systems, analytical databases, large-scale storage backends). Examples of these include Kafka, Redpanda, WarpStream, Postgres, ClickHouse, Snowflake, or Cassandra. • Strong Kubernetes experience in AWS, GCP, or Azure, and familiarity with infrastructure-as-code tooling (Helm, Terraform, Jsonnet, etc.). • Solid understanding of distributed systems design and large-scale system trade-offs. • Proficiency in at least one programming language (Go preferred, but not required). • Working knowledge of Linux internals, networking, cloud storage, and performance/scaling behavior. • Experience participating in blameless incident response and writing high-quality post-incident reviews. • Clear communicator who can collaborate across teams and work autonomously. • Curious, pragmatic, action-oriented, and kind (this is important!)
• 100% Remote, Global Culture • Scaling Organization – Tackle meaningful work in a high-growth, ever-evolving environment. • Transparent Communication – Expect open decision-making and regular company-wide updates. • Innovation-Driven – Autonomy and support to ship great work and try new things. • Open Source Roots – Built on community-driven values that shape how we work. • Empowered Teams – High trust, low ego culture that values outcomes over optics. • Career Growth Pathways – Defined opportunities to grow and develop your career. • Approachable Leadership – Transparent execs who are involved, visible, and human. • Passionate People – Join a team of smart, supportive folks who care deeply about what they do. • In-Person onboarding - We want you to thrive from day 1 with your fellow new ‘Grafanistas’ to learn all about what we do and how we do it. • Balance is Key - We operate a global annual leave policy of 30 days per annum. 3 days of your annual leave entitlement are reserved for Grafana Shutdown Days to allow the team to really disconnect.
Apply Now🔥 3 hours ago
Senior Fullstack Developer enhancing digital communication through Chatbot solutions, working collaboratively in a motivated team in Würzburg. Flexible remote work options available with a focus on software quality and development.
🗣️🇩🇪 German Required
🔥 13 hours ago
Senior Product Engineer owning features from concept to delivery in an AI focus at Journee. Collaborating across teams while working full-stack with React, TypeScript, and Python.
JavaScript
Node.js
Python
React
TypeScript
🔥 20 hours ago
Software Developer for interface & driver communication at Dedalus, enhancing healthcare technology. Develop, integrate, and test systems while ensuring software quality and compliance.
🗣️🇩🇪 German Required
Java
TCP/IP
🕒 Yesterday
51 - 200
Senior Software Engineer enhancing Exporter apps' code quality using AI tools. Collaborating on automated testing and 3rd level support with Atlassian products.
JavaScript
React
Redux
TypeScript
🕒 Yesterday
Lead Engineer managing payment solutions at vivenu, the global leader in event ticketing tech. Scaling payment platforms while ensuring reliability and technical excellence in a rapidly growing company.
AWS
Cloud
Docker
Google Cloud Platform
Kubernetes
NoSQL
SQL
TypeScript
Go