Senior DevOps Engineer

🕒 July 31

🌐 Argentina, Poland, +2 more countries – Remote

infoinfo

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

đŸ‘» Ghost score 10%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Solvd, Inc.

Solvd, Inc.

501 - 1000 employees

Founded 2010

đŸ’Œ Consulting

đŸ„ Healthcare

📩 Logistics

Consulting ‱ Healthcare ‱ Logistics

Solvd, Inc. is a global software engineering and consulting company that delivers comprehensive end-to-end solutions. Founded in 2011, Solvd offers core engineering services, including software development for web and mobile platforms, digital experience and design, data and AI/ML solutions, and cloud platform modernization. The company is an AWS partner and is known for optimizing resources, solving complex problems, and enhancing user satisfaction. Solvd emphasizes innovation and quality assurance, employing over 800 international professionals across 8 offices worldwide. The company prides itself on transforming businesses by providing feature-rich digital products, manual testing, automation, and employing cutting-edge technology solutions. Their proprietary tools like Zebrunner and Carina aid in improving development and QA processes. Solvd operates with a focus on collaboration, top talent cultivation, and customized technical solutions tailored to individual client needs, serving clients across 15 countries.

📋 Description

‱ Design, implement, and maintain secure, reliable, and scalable cloud infrastructure. ‱ Build and improve CI/CD pipelines that support frequent, repeatable, and low-risk deployments. ‱ Automate infrastructure provisioning, configuration, deployment, and environment management. ‱ Maintain infrastructure as code using tools such as Terraform, CloudFormation, Pulumi, or comparable technologies. ‱ Improve consistency across development, testing, staging, and production environments. ‱ Build and maintain containerized application environments and orchestration capabilities. ‱ Implement and improve monitoring, logging, tracing, alerting, dashboards, and production-health reporting. ‱ Partner with engineers to establish SLIs, SLOs, and actionable operational metrics. ‱ Improve application and infrastructure resiliency, scalability, availability, and disaster recovery readiness. ‱ Manage secrets, certificates, identity, permissions, network controls, and cloud security configurations. ‱ Support vulnerability management, dependency security, patching, and infrastructure hardening. ‱ Improve deployment strategies including automated rollback, blue-green deployment, canary releases. ‱ Diagnose and resolve production incidents involving infrastructure, networking, application deployment, performance, capacity, or cloud services. ‱ Participate in incident response, post-incident reviews, root-cause analysis, and corrective-action planning. ‱ Reduce cloud waste and improve infrastructure cost visibility. ‱ Build self-service tools and reusable deployment patterns for application engineering teams. ‱ Document infrastructure architecture, operational procedures, recovery processes, and troubleshooting guidance. ‱ Participate in production support and an appropriate on-call rotation. ‱ Use AI-assisted engineering tools to accelerate scripting, troubleshooting, documentation, and infrastructure analysis.

🎯 Requirements

‱ Approximately 5+ years of experience in DevOps, SRE, Cloud Engineering, Platform Engineering, or closely related role. ‱ Strong experience operating production workloads in AWS. ‱ Strong hands-on experience with infrastructure as code (Terraform, CloudFormation, Pulumi). ‱ Experience building and maintaining CI/CD pipelines (GitHub Actions, GitLab CI, Jenkins, CircleCI, AWS CodePipeline). ‱ Strong knowledge of Docker and containerized application delivery. ‱ Experience with Kubernetes, Amazon ECS, or another container orchestration platform. ‱ Strong understanding of cloud networking (VPCs, subnets, routing, load balancers, DNS, firewalls, gateways, private connectivity). ‱ Experience with cloud IAM, role-based access, least-privilege design, and secrets management. ‱ Experience implementing logging, metrics, dashboards, alerts, and distributed tracing. ‱ Familiarity with observability platforms (Datadog, New Relic, Grafana, Prometheus, CloudWatch, OpenTelemetry). ‱ Strong Linux and command-line skills. ‱ Scripting experience with Python, Bash, JavaScript, or comparable language. ‱ Understanding of application deployment patterns for modern web applications, APIs, background workers, and event-driven systems. ‱ Experience supporting relational databases, backups, restore processes, high availability. ‱ Knowledge of incident-management practices, root-cause analysis, capacity planning, and production-readiness reviews. ‱ Working knowledge of cloud security, vulnerability remediation, encryption, certificates, key management, and compliance-oriented controls.

đŸ–ïž Benefits

‱ Be part of a global team with equal opportunities for collaboration across continents and cultures. ‱ Thrive in an inclusive environment that prioritizes continuous learning, innovation, and ethical AI standards.

Apply Now

Similar Jobs

🕒 July 28

Domus Global

501 - 1000

đŸ’Œ Consulting

📣 Marketing

📩 Logistics

DevOps Engineer designing and optimizing AWS infrastructure and CI/CD processes at Nublit. Responsible for automation and continuous improvement in development teams.

đŸ—ŁïžđŸ‡Ș🇾 Spanish Required

AWS

Cloud

Grafana

Jenkins

Kubernetes

Prometheus

Terraform

🕒 July 27

Cognativ

11 - 50

đŸ’Œ Consulting

đŸ„œ AR/VR

đŸ€– Artificial Intelligence

Senior Site Reliability Engineer ensuring stability of distributed AI alerting platform. Leading incident response, capacity planning, and observability with a focus on service reliability.

Apache

AWS

Cloud

Grafana

IoT

Java

Kafka

Linux

Postgres

Prometheus

Python

Redis

Terraform

Go

🕒 July 24

Software Mind

1001 - 5000

đŸ€– Artificial Intelligence

☁ SaaS

📡 Telecommunications

DevOps Engineer designing Infrastructure as Code and automating deployment processes for a multicultural engineering team at Software Mind. Collaborating on CI/CD pipelines while ensuring system observability and security compliance.

AWS

Azure

Cloud

Docker

Grafana

Kubernetes

Prometheus

Python

Terraform

🕒 July 23

Sezzle

201 - 500

💳 Fintech

đŸ‘„ B2C

đŸ›ïž eCommerce

Senior Site Reliability Engineer at Sezzle enhancing infrastructure reliability and scalability through modern tech solutions. Collaborating across teams while driving innovation in the fintech space.

AWS

Distributed Systems

Grafana

Kubernetes

Microservices

MySQL

Postgres

Prometheus

RDBMS

SQL

Go

🕒 July 14

SecurityScorecard

501 - 1000

đŸ’Œ Consulting

đŸ„ Healthcare

đŸ›Ąïž Insurance

Senior Site Reliability Engineer driving Kubernetes infrastructure design and optimization for SecurityScorecard. Focused on AI tooling, CI/CD systems, and team collaboration.

Grafana

Jenkins

Kafka

Kubernetes

Prometheus

Python

Terraform

Go