Senior DevOps Engineer

🕒 May 13

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Shuru

Shuru

51 - 200 employees

Founded 2021

🤖 Artificial Intelligence

🤝 B2B

🏢 Enterprise

Artificial Intelligence • B2B • Enterprise

Shuru is a product, AI, and technology consulting firm that partners with businesses to deliver strategic consulting, full-cycle product and custom software development, and curated engineering team extension. Their AI-native engineering teams build scalable AI applications, data engineering and analytics, cloud/DevOps, and API integrations to modernize systems and accelerate product delivery. Shuru operates globally with a remote-first model and emphasizes high ownership, design thinking, and measurable outcomes for enterprise and startup clients.

📋 Description

• Help take cloud platform from pre-production to production readiness and scale. • Work closely with engineering and data teams to bring infrastructure under code. • Improve deployment pipelines, set up monitoring and alerting, and support production deployment of data pipelines and risk oracle workloads. • Assess and harden current platform setup primarily on GCP. • Bring infrastructure under Infrastructure-as-Code using Terraform or similar. • Standardize development, staging, and production environments. • Design and operate the platform's networking layer: VPC architecture, private connectivity, load balancing. • Lead decisions on workload management between Cloud Run and GKE. • Harden GitHub Actions pipelines. • Set up monitoring, logging, tracing, alerting, and dashboards across the platform. • Work with data teams to productionize data pipelines and risk oracle workloads. • Establish secrets management, audit logging, IAM, and access patterns.

🎯 Requirements

• 5+ years of DevOps, Platform, or SRE experience, ideally with at least 2 years working on GCP; experience with vertex ai, AWS or Azure is a plus. • Hands-on production experience with Infrastructure-as-Code tools such as Terraform, Pulumi, CDK, or similar. • Strong CI/CD experience, especially with GitHub Actions or similar. • Experience deploying and operating containerized services using Cloud Run, Kubernetes/GKE, ECS, or similar platforms. • Good judgment on when to use managed or serverless platforms versus Kubernetes or orchestrated approaches. • Experience operating production data and caching infrastructure, including Cloud SQL/Postgres, Redis/Memorystore. • Experience setting up production monitoring, logging, alerting, dashboards, and reliability targets. • Solid understanding of cloud security fundamentals, including IAM, secrets management, audit logging. • Experience with workflow orchestration or async task systems such as Temporal, Celery or similar. • Experience supporting ML or AI inference workloads in production, with hands-on experience across vector databases.

🏖️ Benefits

• Work on global projects with clients from worldwide. • Be part of a remote-first culture-work from anywhere with flexibility. • Enjoy team-building activities and regular outings. • Collaborate and grow in a supportive environment with opportunities to learn from senior engineers. • Competitive salary and benefits package.

Apply Now

Similar Jobs

🕒 May 12

Volvo Cars

10,000+ employees

🏭 Manufacturing

🚗 Transport

🚘 Automotive

Salesforce Release Engineer driving digital innovation at Volvo Cars. Managing Salesforce release lifecycle across global teams and developing cutting-edge technology solutions for the automotive industry.

🕒 May 12

Akamai Technologies

5001 - 10000

🔒 Cybersecurity

Site Reliability Engineer II managing cloud platforms and enabling advanced computing for customers at Akamai. Collaborating with teams to maintain SLO-driven reliability and secure services.

Ansible

Chef

Distributed Systems

DNS

Grafana

Jenkins

Linux

Prometheus

Python

SaltStack

Shell Scripting

Terraform

🕒 May 5

AlphaSense

1001 - 5000

💼 Consulting

🏥 Healthcare

📣 Marketing

Cloud Reliability & Recovery Engineer at AlphaSense designing and implementing disaster recovery capabilities across AWS. Collaborating with cross-functional teams on high-availability cloud architectures.

AWS

Cloud

DNS

DynamoDB

EC2

Kubernetes

Python

Terraform

🕒 May 5

AlphaSense

1001 - 5000

💼 Consulting

🏥 Healthcare

📣 Marketing

Cloud Reliability & Recovery Engineer focusing on designing and improving AWS BCP and DR capabilities at AlphaSense, a market intelligence company. Collaborates across teams for system resilience and recovery from disruptions.

AWS

Cloud

DNS

DynamoDB

EC2

Kubernetes

Python

Terraform

🕒 April 29

Tookitaki

51 - 200

🤖 Artificial Intelligence

Site Reliability Engineer maintaining and scaling infrastructure for fintech solutions at Tookitaki. Collaborating with engineering and DevOps teams for high availability and performance.

Ansible

AWS

Cloud

Docker

Google Cloud Platform

Grafana

Kubernetes

Linux

MariaDB

Prometheus

Python

SQL

Terraform