Senior Software Engineer – Site Reliability

🔥 14 hours ago

🇺🇸 United States – Remote

💵 $114.8k - $150k / year

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

👻 Ghost score 0%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Bloomerang

Bloomerang

201 - 500 employees

Founded 2012

💼 Consulting

📣 Marketing

🤝 Non-profit

💰 $33M Debt Financing on 2021-02

Consulting • Marketing • Non-profit

Bloomerang is a community-focused nonprofit donor management software designed to deliver a better giving experience and help organizations thrive. It offers a comprehensive platform tailored for nonprofits, including features such as donor database management, marketing and engagement tools, reporting and analytics solutions, and volunteer and membership management systems. Bloomerang aims to enhance donor retention and increase fundraising efficacy by equipping nonprofits with the tools they need to nurture relationships and drive donations. Trusted by over 23,000 organizations, Bloomerang helps nonprofits raise more funds and create lasting change through streamlined operations and an intuitive user experience.

📋 Description

• Own complex production support escalations and ticket triage, providing hands-on troubleshooting and resolution • Partner with Software Engineering to investigate production issues, identify root causes and reliability risks, and drive permanent solutions • Apply and mature SRE practices, proactive reliability engineering, continuous improvement, automation, and shared ownership • Lead incident response from triage and mitigation through recovery, root cause analysis, and blameless post-incident reviews • Build observability using metrics, logs, traces, dashboards, and actionable alerts • Define and mature SLIs and SLOs measuring system reliability and customer experience • Develop synthetic monitoring for critical customer journeys • Identify operational toil and drive automation, tooling, process improvements, or permanent fixes • Use AI-assisted tools and source code repositories for triage, troubleshooting, code analysis, automation, and technical investigation • Participate in a rotating on-call schedule, primarily during business hours, with limited after-hours and weekend support

🎯 Requirements

• Hands-on Site Reliability Engineering experience applying software engineering practices to production reliability and helping establish or mature SRE practices • Strong knowledge of SLIs, SLOs, error budgets, observability, automation, and toil reduction • Experience building monitoring, dashboards, alerts, and telemetry using Honeycomb, New Relic, Grafana, CloudWatch, Kibana, or similar • Experience with production incident management, root cause analysis, blameless post-incident reviews, and corrective-action follow-through • Strong programming and scripting skills • Strong SQL and relational database skills; PostgreSQL experience preferred • Strong code literacy and debugging skills, including navigating unfamiliar codebases, understanding application flow, reviewing code and change history, and identifying reliability issues • Experience troubleshooting cloud-hosted applications using source code, logs, APIs, telemetry, event streams, and databases • Comfort navigating application stacks across PHP, .NET, and Node.js; deep expertise in each is not required • Demonstrated experience using AI-assisted tools in engineering workflows • Collaborative self-starter able to tackle difficult problems, adapt to changing priorities, and challenge the status quo • Strong communication and collaboration skills across Software Engineering, Product, Support, DevOps, and other technical teams • Ability to work within the U.S. and select Canadian Provinces • No visa sponsorship available

🏖️ Benefits

• Health, vision, and dental insurance options • HealthiestYou healthcare service with confidential access to doctors 24/7 • 20 PTO days • 3 flex days • 4 optional volunteer days • 12 paid holidays • Paid parental leave • 401k match • Equipment shipped to your door • Potential eligibility for a discretionary bonus • Fully remote work within the U.S. and select Canadian Provinces

Apply Now

Similar Jobs

🔥 15 hours ago

High 5 Games

51 - 200

🎮 Gaming

🎲 Gambling

🤝 B2B

DevOps Engineer scaling Google Cloud infrastructure for machine learning operations. Automating deployments, pipelines, monitoring, and reliability for AI systems serving millions of players.

🔥 21 hours ago

CoorsTek, Inc.

5001 - 10000

🏭 Manufacturing

🚘 Automotive

🎖️ Defense

AI Deployment Engineer delivering secure production applications and AI automation for CoorsTek’s advanced ceramics manufacturing operations. Translating plant and business needs into scalable software.

🔥 21 hours ago

Particle41

51 - 200

💼 Consulting

📣 Marketing

☁️ SaaS

Lead DevOps Engineer owning cloud infrastructure, automation, reliability, and security. Advising Particle41 clients and mentoring teams delivering complex digital projects.

🔥 23 hours ago

Astronomer

201 - 500

🏥 Healthcare

🏭 Manufacturing

💼 Consulting

Customer Reliability Engineer supporting Astronomer’s managed Apache Airflow platform. Troubleshooting customer environments across Python, Airflow, Kubernetes, containers, and cloud systems.

🕒 Yesterday

9th Way Insignia

51 - 200

💼 Consulting

🏥 Healthcare

📦 Logistics

DevSecOps Engineer securing VA cloud infrastructure, applications, and CI/CD pipelines for government cybersecurity modernization. Automating compliance, vulnerability management, software delivery, and Authority to Operate processes.