Staff DevOps Engineer

🔥 2 minutes ago

🇪🇸 Spain – Remote

⏰ Full Time

🔴 Lead

⛑ DevOps & Site Reliability Engineer (SRE)

👻 Ghost score 21%

infoinfo

🗣️🇷🇺 Russian Required

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of MEDvidi

MEDvidi

201 - 500 employees

Founded 2019

🏥 Healthcare

⚕️ Healthcare Insurance

💰 $2.8M Seed Round on 2022-09

Healthcare • Healthcare Insurance • Mental Health

MEDvidi is an online mental health treatment center aiming to make professional care accessible and affordable for everyone. The healthcare experts at MEDvidi provide personalized treatment plans for a variety of mental health conditions, including ADHD, anxiety, depression, insomnia, and OCD. With services that include initial assessments, ongoing support, and medication management through virtual consultations, MEDvidi seeks to enhance mental wellness with compassionate care tailored to individual needs.

📋 Description

• Set and drive the technical vision and quarterly roadmap for infrastructure with clear trade-offs and measurable goals • Run and evolve AWS and Kubernetes/EKS infrastructure, including cluster management, Karpenter autoscaling, Kyverno policy enforcement, and zero-downtime operations • Own Infrastructure as Code with Terraform and AWS CDK in TypeScript • Own GitLab CI/CD, including reusable/shared templates, OIDC, and self-managed GitLab • Build and own practical observability using Prometheus, Grafana, OpenTelemetry, OpenSearch, CloudWatch, log pipelines, and APM • Maintain fast blue-green deployments and health-gated automated rollback • Own zero-downtime PostgreSQL schema migrations and CI migration gating • Own security engineering in a HIPAA environment, including secrets hygiene, credential rotation, leak scanning, PHI-aware log and data handling, and Vault managed as code • Partner with product engineering teams to remove infrastructure friction and improve developer experience • Use agentic AI as a core workflow, integrating autonomous-agent output into production • Set technical direction, ship infrastructure hands-on, and own reliability, cost, security, performance, deployment health, and developer velocity outcomes

🎯 Requirements

• 6+ years in DevOps/infrastructure engineering • Strong systems fundamentals and solid Linux administration and troubleshooting, including performance analysis, resource management, and process debugging • Hands-on AWS experience with EC2, EKS, RDS, ElastiCache, Lambda, SQS, EventBridge, API Gateway, ALB, and S3 • Production Kubernetes/EKS experience, including cluster management, node scaling, and policy enforcement; Karpenter, Kyverno, or similar • Strong Infrastructure as Code experience with Terraform and AWS CDK in TypeScript • CI/CD ownership with GitLab CI/CD, reusable/shared templates, OIDC id_tokens, and self-managed GitLab • Practical monitoring and observability experience with Prometheus, Grafana, OpenTelemetry, OpenSearch, CloudWatch, log-shipping, and error tracking/APM • Practical security engineering experience with secrets rotation, short-lived credentials, leak scanning, and PHI-aware logging • HashiCorp Vault as code experience, including KV, JWT/OIDC authentication for CI, and policy design • Experience with blue-green deployments and automated, health-gated rollback • Experience with PostgreSQL zero-downtime schema migrations using expand/contract and migration gating in CI • Container experience with Docker, ECR, immutable tags, and image lifecycle management • Network/protocol fundamentals including load balancing, TLS, and DNS • Hands-on agentic AI workflows, such as Claude Code or similar • Developer-focused mindset, strong problem-solving for complex system issues, and strong technical writing, including docs-as-code, ADRs, and design docs via MRs • Fluent Russian and English (B2) • Experience working effectively in remote, distributed teams • Experience in a regulated/compliance-heavy environment, such as HIPAA or SOC 2, is a plus • Configuration management with Ansible is a plus • Node.js application operations with pm2 and npm is a plus • GitOps tooling such as ArgoCD or Flux and deeper PostgreSQL database administration are a plus • AWS certifications are a plus

🏖️ Benefits

• Competitive compensation package • Health insurance after the probation period • Sports & wellness compensation • Personalized English lessons via Preply • 19 paid vacation days annually • 4 additional wellness days each year • Paid sick leave for the first 5 working days • Thoughtful gifts for key life events • Offline corporate events • Fully remote long-term collaboration under a B2B model

Apply Now

Similar Jobs

🕒 May 12

Stellar Cyber

51 - 200

🔒 Cybersecurity

🤖 Artificial Intelligence

🏢 Enterprise

Seeking a Staff Site Reliability Engineer to enhance scalability and efficiency for production systems at Stellar Cyber. Join a leader in cybersecurity focused on innovative solutions against cyber threats.

AWS

Azure

Cloud

Distributed Systems

ElasticSearch

Google Cloud Platform

Grafana

Kafka

Kubernetes

Linux

MongoDB

Prometheus

Python

Redis

Spark

Terraform