Cloud Systems Engineer – Site Reliability

🔥 1 hour ago

🔔 Pennsylvania – Remote

infoinfo

💵 $110k - $150k / year

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

👻 Ghost score 0%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of TherapyNotes, LLC

TherapyNotes, LLC

51 - 200 employees

Founded 2010

💼 Consulting

⚖️ Legal

🏥 Healthcare

Consulting • Legal • Healthcare

TherapyNotes, LLC is a comprehensive practice management system designed specifically for behavioral health professionals, including psychologists, therapists, psychiatrists, and social workers. The platform offers an array of features such as scheduling, telehealth, electronic health records (EHR), billing, and a client portal, all integrated into a secure and user-friendly software solution. TherapyNotes aims to streamline clinical workflows, enhance patient care, and reduce administrative burdens for mental health practices, supporting them with dedicated customer service and continuous innovation based on user feedback.

📋 Description

• Own and continuously improve use of Datadog across metrics, logs, traces, dashboards, monitors, alerts, and service-level views • Design, implement, and maintain high-availability, high-throughput, data- and compute-intensive critical systems for a 24×7 SaaS platform • Partner with service owners to define and improve reliability through SLIs, SLOs, error budgets, actionable alerting, and operational-readiness practices • Participate in incident management as incident commander or technical responder, coordinating triage, restoration, escalation, communication, documentation, root cause analysis, and corrective actions • Investigate issues across infrastructure and application layers using metrics, logs, distributed traces, and code-level context • Improve deployment safety and service resilience through automated validation, recovery and rollback capabilities, reliability testing, and failure-mode analysis • Ensure newly introduced systems are supportable and maintainable by development and operations • Provide escalated technical guidance and support to technology teams • Provide on-call coverage for production support and other duties as required • Ensure systems and operational activities comply with organizational security, HIPAA, and operating policies • Eliminate repetitive operational toil using Bash, PowerShell, Python, or Ansible • Manage infrastructure as code using Terraform/OpenTofu and configuration automation using Ansible

🎯 Requirements

• BS degree in Information Systems, Engineering, or equivalent experience • 5+ years of engineering experience in Systems Engineering, Cloud or Platform Engineering, DevOps, Software Engineering, and/or SRE • Experience designing and operating production systems using cloud-based compute, storage, networking, and containerization technologies; Azure and Kubernetes preferred • Strong Linux systems and networking fundamentals, with experience troubleshooting complex distributed systems in production • Expertise with an observability platform; Datadog experience strongly preferred • Experience with Prometheus, Grafana, New Relic, or equivalent platforms is valuable • Experience with scripting and operational automation using tools such as Bash, PowerShell, or Python, along with infrastructure-as-code and configuration-management practices • Experience participating in production on-call rotations, incident response, root cause analysis, and post-incident improvement • Experience working in Agile/DevOps environments and operating production services using ITSM practices where applicable • Prior software development experience—or experience investigating application behavior through code, logs, and distributed traces—is a plus

🏖️ Benefits

• Employer sponsored health, dental, vision, life, and disability insurance • Retirement plan with company contribution • Annual company profit sharing • Personal development/training budget • Open, collaborative work environment • Extensive 2-week onboarding plan • Comprehensive mentorship program

Apply Now

Similar Jobs

🔥 7 hours ago

Sphera

1001 - 5000

💼 Consulting

🏥 Healthcare

📦 Logistics

Cybersecurity Engineer securing Sphera’s environmental, health, safety, and sustainability software for U.S. government and DoD clients. Managing RMF, STIG, ATO, vulnerability remediation, and DevSecOps compliance.

🔥 9 hours ago

Arista Networks

1001 - 5000

🏢 Enterprise

📡 Telecommunications

FedRAMP SRE operating Arista Networks’ Kubernetes-native CloudVision networking SaaS. Ensuring reliable, secure, scalable production systems and leading infrastructure projects.

🔥 9 hours ago

Gainwell Technologies

10,000+ employees

💼 Consulting

📦 Logistics

⚕️ Healthcare Insurance

DevOps Engineer automating configuration, releases, and cloud infrastructure for Gainwell Technologies’ healthcare platforms. Building scripts, CI processes, and infrastructure as code for SaaS and PaaS products.

🔥 9 hours ago

Gainwell Technologies

10,000+ employees

💼 Consulting

📦 Logistics

⚕️ Healthcare Insurance

DevOps Engineer automating infrastructure, configuration management, and software releases. Supporting Gainwell’s cloud-based healthcare technology platforms and development teams.

🔥 11 hours ago

Coalfire

1001 - 5000

💼 Consulting

🏥 Healthcare

📦 Logistics

Senior Site Reliability Engineer operating FedRAMP-compliant cloud environments for Coalfire, a cybersecurity consulting firm. Owning observability, automation, incident response, recovery, and compliance operations.