Site Reliability Engineer

🕒 June 27

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Berkeley Research Group (BRG)

Berkeley Research Group (BRG)

1001 - 5000 employees

🏗️ Construction

📦 Logistics

🎖️ Defense

💰 Venture Round on 2020-07

Construction • Logistics • Defense

Berkeley Research Group (BRG) is a global consulting firm that helps leading organizations advance in the fields of corporate finance; economics, disputes, and investigations; and performance improvement. With offices around the world, BRG is an integrated group of experts, industry leaders, academics, data scientists, and professionals working across borders and disciplines. The firm specializes in various sectors, including construction, energy, technology, and healthcare, delivering inspired insights and practical strategies to help clients navigate their challenges effectively.

📋 Description

• Design, implement, and maintain scalable and reliable systems in cloud environments such as Azure Cloud Services. • Provide operational support for full-stack software applications. • Increase system resilience with expert-level coding, bulletproof release, and change management skills. • Develop service-level indicators and objectives to automate release validation. • Improve automation and increase the system’s self-healing capability. • Collect operating system data and report performance metrics to stakeholders. • Ensure security best practices are followed in cloud infrastructure and application deployments. • Manage cloud and database system maintenance, debugging production issues as they arise. • Improve reliability, quality, and time-to-market of our suite of software solutions. • Partner with security and product teams to define and publish policies, processes, and playbooks to facilitate rapid and effective handling of alerts and incidents. • Lead incident management processes; respond to outages and service disruptions promptly.

🎯 Requirements

• Bachelor’s degree in computer science or similar field. • Five years’ experience as a site reliability engineer or similar role. • Strong programming skills (Golang, Ruby, Python, or similar). • Proven ability to diagnose and monitor performance and reliability issues across the stack. • Expertise in Kubernetes. • Relevant industry certifications, such as through the Site Reliability Engineering (SRE) Foundation. • Proven experience working with cloud-native infrastructure (Azure Cloud Services, AWS, or GCP). • Experience working with observability and incident management tools (Datadog, OpsGenie, PagerDuty). • Experience scripting operating system tasks with Infrastructure as Code. • Impeccable communication skills. • Ability to problem-solve in a fast-paced, high-stakes environment. • Candidate must be able to submit verification of his/her legal right to work in the United States, without company sponsorship.

Apply Now

Similar Jobs

🕒 June 26

CACI International Inc

10,000+ employees

💼 Consulting

🎖️ Defense

Cloud DevOps Engineer managing CI/CD pipelines and applications in AWS cloud. Collaborating on security initiatives and providing DevSecOps training with Agile teams.

🇺🇸 United States – Remote

💵 $82.1k - $172.4k / year

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🕒 June 26

Alkami Technology

501 - 1000

💼 Consulting

📣 Marketing

🏦 Banking

Site Reliability Engineer at Alkami developing and testing code for application releases. Collaborating with teams to improve delivery and participate in on-call rotations.

🇺🇸 United States – Remote

💵 $110k - $137.5k / year

💰 $300M Post-IPO Debt - Alkami Technology on 2025-03

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🕒 June 26

Hearst Health

1001 - 5000

🏥 Healthcare

⚕️ Healthcare Insurance

☁️ SaaS

Senior DevOps Engineer helping build and improve healthcare technology platforms at Zynx Health. Operating cloud infrastructure, deployment pipelines, and security practices for clinical decision support solutions.

🇺🇸 United States – Remote

💵 $112k - $135k / year

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🕒 June 26

THEMIS Waste Recovery Technology

11 - 50

🏥 Healthcare

📦 Logistics

🍽️ Food & Beverage

DevOps Engineer managing cloud infrastructure and CI/CD pipelines for secure, reliable operations at Themis. Ensuring system health and performance while automating processes and maintaining security controls.

🇺🇸 United States – Remote

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🕒 June 26

NBA

11 - 50

💼 Consulting

🏗️ Construction

🏠 Real Estate

Senior Site Reliability Engineer ensuring the operational excellence of NBA's messaging platforms. Supporting critical operations including live games and global events with expertise in Microsoft Exchange and collaboration services.

SMTP