Senior DevOps Engineer

🔥 7 minutes ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Solvd, Inc.

Solvd, Inc.

501 - 1000 employees

Founded 2010

💼 Consulting

🏥 Healthcare

📦 Logistics

Consulting • Healthcare • Logistics

Solvd, Inc. is a global software engineering and consulting company that delivers comprehensive end-to-end solutions. Founded in 2011, Solvd offers core engineering services, including software development for web and mobile platforms, digital experience and design, data and AI/ML solutions, and cloud platform modernization. The company is an AWS partner and is known for optimizing resources, solving complex problems, and enhancing user satisfaction. Solvd emphasizes innovation and quality assurance, employing over 800 international professionals across 8 offices worldwide. The company prides itself on transforming businesses by providing feature-rich digital products, manual testing, automation, and employing cutting-edge technology solutions. Their proprietary tools like Zebrunner and Carina aid in improving development and QA processes. Solvd operates with a focus on collaboration, top talent cultivation, and customized technical solutions tailored to individual client needs, serving clients across 15 countries.

📋 Description

• Design, implement, and maintain secure, reliable, and scalable cloud infrastructure. • Build and improve CI/CD pipelines that support frequent, repeatable, and low-risk deployments. • Automate infrastructure provisioning, configuration, deployment, and environment management. • Maintain infrastructure as code using tools such as Terraform, CloudFormation, Pulumi, or comparable technologies. • Improve consistency across development, testing, staging, and production environments. • Build and maintain containerized application environments and orchestration capabilities. • Implement and improve monitoring, logging, tracing, alerting, dashboards, and production-health reporting. • Partner with engineers to establish SLIs, SLOs, and actionable operational metrics. • Improve application and infrastructure resiliency, scalability, availability, and disaster recovery readiness. • Manage secrets, certificates, identity, permissions, network controls, and cloud security configurations. • Support vulnerability management, dependency security, patching, and infrastructure hardening. • Improve deployment strategies including automated rollback, blue-green deployment, canary releases. • Diagnose and resolve production incidents involving infrastructure, networking, application deployment, performance, capacity, or cloud services. • Participate in incident response, post-incident reviews, root-cause analysis, and corrective-action planning. • Reduce cloud waste and improve infrastructure cost visibility. • Build self-service tools and reusable deployment patterns for application engineering teams. • Document infrastructure architecture, operational procedures, recovery processes, and troubleshooting guidance. • Participate in production support and an appropriate on-call rotation. • Use AI-assisted engineering tools to accelerate scripting, troubleshooting, documentation, and infrastructure analysis.

🎯 Requirements

• Approximately 5+ years of experience in DevOps, SRE, Cloud Engineering, Platform Engineering, or closely related role. • Strong experience operating production workloads in AWS. • Strong hands-on experience with infrastructure as code (Terraform, CloudFormation, Pulumi). • Experience building and maintaining CI/CD pipelines (GitHub Actions, GitLab CI, Jenkins, CircleCI, AWS CodePipeline). • Strong knowledge of Docker and containerized application delivery. • Experience with Kubernetes, Amazon ECS, or another container orchestration platform. • Strong understanding of cloud networking (VPCs, subnets, routing, load balancers, DNS, firewalls, gateways, private connectivity). • Experience with cloud IAM, role-based access, least-privilege design, and secrets management. • Experience implementing logging, metrics, dashboards, alerts, and distributed tracing. • Familiarity with observability platforms (Datadog, New Relic, Grafana, Prometheus, CloudWatch, OpenTelemetry). • Strong Linux and command-line skills. • Scripting experience with Python, Bash, JavaScript, or comparable language. • Understanding of application deployment patterns for modern web applications, APIs, background workers, and event-driven systems. • Experience supporting relational databases, backups, restore processes, high availability. • Knowledge of incident-management practices, root-cause analysis, capacity planning, and production-readiness reviews. • Working knowledge of cloud security, vulnerability remediation, encryption, certificates, key management, and compliance-oriented controls.

🏖️ Benefits

• Be part of a global team with equal opportunities for collaboration across continents and cultures. • Thrive in an inclusive environment that prioritizes continuous learning, innovation, and ethical AI standards.

Apply Now

Similar Jobs

🕒 Yesterday

intive

1001 - 5000

💼 Consulting

🏥 Healthcare

📣 Marketing

DevOps / Cloud Engineer supporting Azure infrastructure and cloud operations at intive. Managing cloud infrastructure, CI/CD pipelines, and ensuring best practices in a dynamic environment.

Azure

Cloud

Docker

Kubernetes

MS SQL Server

SQL

Terraform

🕒 2 days ago

Teladoc Health

5001 - 10000

🏥 Healthcare

👥 B2C

☁️ SaaS

Site Reliability Engineer at Teladoc Health specializing in Azure observability and monitoring. Ensuring reliability and performance of hybrid cloud infrastructure and services.

🇦🇷 Argentina – Remote

💰 $80M Post-IPO Debt - Teladoc Health on 2016-07

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

Ansible

Azure

Grafana

Python

Terraform

🕒 2 days ago

Domus Global

501 - 1000

💼 Consulting

📣 Marketing

📦 Logistics

DevOps Engineer designing and optimizing AWS infrastructure and CI/CD processes at Nublit. Responsible for automation and continuous improvement in development teams.

🗣️🇪🇸 Spanish Required

AWS

Cloud

Grafana

Jenkins

Kubernetes

Prometheus

Terraform

🕒 3 days ago

Cognativ

11 - 50

💼 Consulting

🥽 AR/VR

🤖 Artificial Intelligence

Senior Site Reliability Engineer ensuring stability of distributed AI alerting platform. Leading incident response, capacity planning, and observability with a focus on service reliability.

Apache

AWS

Cloud

Grafana

IoT

Java

Kafka

Linux

Postgres

Prometheus

Python

Redis

Terraform

Go

🕒 July 24

Software Mind

1001 - 5000

🤖 Artificial Intelligence

☁️ SaaS

📡 Telecommunications

DevOps Engineer designing Infrastructure as Code and automating deployment processes for a multicultural engineering team at Software Mind. Collaborating on CI/CD pipelines while ensuring system observability and security compliance.

AWS

Azure

Cloud

Docker

Grafana

Kubernetes

Prometheus

Python

Terraform