
501 - 1000 employees
Founded 2014
🏢 Enterprise
☁️ SaaS
🤖 Artificial Intelligence
Enterprise • SaaS • Artificial Intelligence
Grafana Labs is a company that specializes in open-source observability technologies and solutions. It offers a comprehensive suite of tools for logging, metrics, tracing, and profile management with products like Grafana, Loki, Tempo, and Mimir. Their offerings are designed to help businesses visualize, monitor, and alert on data from various sources, providing capabilities such as anomaly detection, root cause analysis, and service level objective management using AI/ML insights. Grafana Labs provides both cloud-based and self-managed solutions, ideal for infrastructure, application, and frontend observability. Additionally, their platform supports integration with various data sources like Prometheus and OpenTelemetry, making them a key player in the observability and infrastructure monitoring space.
🔥 13 hours ago
🇺🇸 United States – Remote
💵 $154.4k - $185.3k / year
⏰ Full Time
🟠 Senior
⛑ DevOps & Site Reliability Engineer (SRE)
🦅 H1B Visa Sponsor
Improve your chances of getting an interview by checking your resume score before you apply.

501 - 1000 employees
Founded 2014
🏢 Enterprise
☁️ SaaS
🤖 Artificial Intelligence
Enterprise • SaaS • Artificial Intelligence
Grafana Labs is a company that specializes in open-source observability technologies and solutions. It offers a comprehensive suite of tools for logging, metrics, tracing, and profile management with products like Grafana, Loki, Tempo, and Mimir. Their offerings are designed to help businesses visualize, monitor, and alert on data from various sources, providing capabilities such as anomaly detection, root cause analysis, and service level objective management using AI/ML insights. Grafana Labs provides both cloud-based and self-managed solutions, ideal for infrastructure, application, and frontend observability. Additionally, their platform supports integration with various data sources like Prometheus and OpenTelemetry, making them a key player in the observability and infrastructure monitoring space.
• Partner closely with product engineering squads (embedded model) • Own production reliability for high-SLA and complex customer environments • Design and implement automation to scale our reliability practices • Ensuring our customers meet our SLO targets • Define and evolve per-tenant SLOs and reliability models • Proactively reduce SLO burn to prevent repeat incidents • Serving as a primary escalation point and on-call for relevant incidents • Lead customer-impacting incident response and post-incident reviews • Contribute to design docs and code reviews • Influence feature design to ensure production scalability and operability • Build automation to eliminate toil where needed • Improve alert quality and reduce noisy escalations
• 6+ years engineering experience, 3+ in SRE/CRE/production engineering. Strong preference for those with formal customer reliability engineering experience. • Strong Kubernetes experience in AWS, GCP, or Azure, and familiarity with infrastructure-as-code tooling (Helm, Terraform, Jsonnet, etc.). • Experience operating multi-tenant systems in production • Strong experience designing and implementing SLOs • Experience with one or more programming languages (e.g. Go, Python, Java, etc) • Experience with Linux operating systems internals, and some knowledge of networking, cloud storage, and scaling. • Excellent problem-solving and troubleshooting skills. • Experience with calmly and actively participating in blame-free Incident Response, following up on actions, and writing high quality PIRs (Post Incident Reviews, a.k.a. post-mortem documents) • Ability to reason about performance, scaling, and failure modes • Comfortable working within an engineering team where individuals are encouraged to have a strong sense of autonomy and self-direction. • Ability to partner deeply with product engineering teams • We highly value those who are intellectually curious, who default to transparency, possess a high bias towards action, and who are also kind (this is important!)
• Restricted Stock Units (RSUs) • 30 days annual leave • Grafana Shutdown Days to allow team to disconnect
Apply Now🔥 13 hours ago
Lead DevOps Engineer implementing CI/CD pipelines and cloud infrastructure at Wolters Kluwer. Mentoring a team and improving system reliability through continuous development practices.
🇺🇸 United States – Remote
💵 $116.4k - $204.1k / year
⏰ Full Time
🟠 Senior
⛑ DevOps & Site Reliability Engineer (SRE)
🦅 H1B Visa Sponsor
Ansible
AWS
Azure
Cloud
Java
Jenkins
Python
Shell Scripting
Terraform
.NET
🔥 14 hours ago
Lead reliability workstreams for Akamai's serverless inference platform. Design SRE tooling and automation while mentoring other SREs in cross-team initiatives.
🇺🇸 United States – Remote
💵 $146.4k - $263.6k / year
💰 Post-IPO Equity on 2001-07
⏰ Full Time
🟠 Senior
⛑ DevOps & Site Reliability Engineer (SRE)
🦅 H1B Visa Sponsor
Distributed Systems
Kubernetes
Python
Go
🔥 16 hours ago
Infrastructure Deployment Engineer overseeing physical installation projects for AI infrastructure. Collaborating with cross-functional teams to ensure quality and compliance in installations.
🇺🇸 United States – Remote
💵 $218k - $263k / year
⏰ Full Time
🟡 Mid-level
🟠 Senior
⛑ DevOps & Site Reliability Engineer (SRE)
🕒 3 days ago
DevOps Engineer IV for Envision's Engineering team. Collaborating on implementing advanced CI/CD pipelines and infrastructure as code.
🇺🇸 United States – Remote
⏰ Full Time
🟡 Mid-level
🟠 Senior
⛑ DevOps & Site Reliability Engineer (SRE)
🦅 H1B Visa Sponsor
Airflow
Amazon Redshift
AWS
Cloud
Docker
ETL
Jenkins
Kubernetes
Python
Spark
SQL
SSIS
Terraform
🕒 3 days ago
Corporate Reliability Engineer enhancing reliability and performance at Arclin's manufacturing processes. Involves data analysis, reliability assessments, and cross-functional collaboration for continuous improvement.