Search Remote Jobs

Senior Software Engineer, DevOps

🔥 33 minutes ago

🇺🇸 United States – Remote

⏰ Full Time

đźź  Senior

⛑ DevOps & Site Reliability Engineer (SRE)

đź‘» Ghost score 25%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Atria Institute

Atria Institute

51 - 200 employees

🏥 Healthcare

⚕️ Healthcare Insurance

🔬 Science

Healthcare • Healthcare Insurance • Science

Atria Institute is leading a movement towards proactive, preventive health care by integrating cutting-edge science and technology into medicine. Comprising the Atria Academy of Science & Medicine, a team of multidisciplinary experts, and the Atria Health Collaborative, a nonprofit organization, Atria seeks to transform health care delivery both locally and globally. The institute offers personalized, coordinated health care and predictive testing, focusing on advanced screening for diseases. Atria is dedicated to translating scientific research into real-time medical applications and implementing proven interventions for preventable diseases.

đź“‹ Description

• Design, build, and maintain Google Cloud Platform infrastructure using Terraform. • Define infrastructure patterns and standards. • Own and improve GitHub Actions CI/CD pipelines. • Identify and eliminate operational toil through scalable automation and tooling. • Establish monitoring, dashboards, and alerting in Datadog, and reduce alert noise. • Lead incident response within the on-call rotation and drive postmortems and follow-ups. • Define and track Service Level Objectives for core infrastructure and build systems. • Partner with product engineering teams on deployment pipelines, environment issues, and build troubleshooting. • Own preview and staging environments, including data sync, masking, and cleanup routines. • Improve developer experience through tooling, runbooks, and documentation. • Write tested and reviewed infrastructure code. • Lead design reviews and RFCs with an operations-focused perspective. • Design systems balancing reliability, performance, and security. • Conduct code reviews and mentor engineers through pairing, knowledge sharing, and documentation. • Collaborate with the Tech Lead, DevOps, Platform, Data Engineering, Clinical Experience, Member Experience, and Care Delivery teams.

🎯 Requirements

• ~5+ years of professional experience in DevOps, SRE, infrastructure, or backend engineering in production environments. • Hands-on experience designing and operating infrastructure in at least one cloud provider, ideally Google Cloud Platform. • Track record of owning and shipping automation, pipelines, or infrastructure that made a team measurably more productive or reliable. • An enthusiasm for developer productivity and making teams as impactful as possible. • Deep experience with infrastructure-as-code, ideally Terraform, and building CI/CD pipelines, ideally GitHub Actions. • Proficient software engineering ability and strong command of Linux. • Strong instincts for monitoring and observability, with confident debugging across logs, traces, and metrics. • Solid experience with relational databases, including MySQL and PostgreSQL, and containerized workloads. • Strong grounding in reliability, performance, and security fundamentals, with judgment to make sound tradeoffs. • Nice to have: experience in healthcare, digital health, or regulated domains such as HIPAA, PHI, and SOC 2. • Nice to have: experience with containers and orchestration, including Docker and Kubernetes. • Nice to have: exposure to leading incident response, on-call, and postmortem practices. • Nice to have: experience with database migrations or managing multiple environments at scale.

Apply Now

Similar Jobs

🔥 2 hours ago

Dental Intelligence

2 - 10

🏥 Healthcare

🏭 Manufacturing

🤝 B2B

Senior SRE scaling Dental Intelligence’s dental-practice SaaS platform across Azure, AWS, and Windows on-premises infrastructure. Leading Terraform, security, observability, and reliability modernization.

🔥 12 hours ago

Experian

10,000+ employees

đź’Ľ Consulting

📣 Marketing

📦 Logistics

SRE plena operando plataformas AWS, Kubernetes e dados para a Experian, empresa global de dados e tecnologia. Melhorando confiabilidade, observabilidade, automação e custos de sistemas cloud-native.

🗣️🇧🇷🇵🇹 Portuguese Required

🔥 12 hours ago

Bitdeer Group

201 - 500

đź’Ľ Consulting

📦 Logistics

🏗️ Construction

Senior storage SRE operating high-performance systems for Bitdeer's AI GPU cloud. Designing resilient storage, telemetry, and automation for large-scale training clusters.

🔥 12 hours ago

Bitdeer Group

201 - 500

đź’Ľ Consulting

📦 Logistics

🏗️ Construction

Senior Kubernetes SRE operating GPU cloud control planes for Bitdeer’s AI and Bitcoin mining infrastructure. Automating scheduling, multi-tenant isolation, BMaaS, and AIOps remediation.

🔥 12 hours ago

Bitdeer Group

201 - 500

đź’Ľ Consulting

📦 Logistics

🏗️ Construction

Senior network SRE owning EVPN/BGP infrastructure and automation for Bitdeer's global AI GPU cloud. Connecting US, APAC, and Iceland data centers.