Staff DevOps Engineer

Job not on LinkedIn

đŸ”„ 8 minutes ago

đŸ—ŁïžđŸ‡·đŸ‡ș Russian Required

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of MEDvidi

MEDvidi

201 - 500 employees

Founded 2019

đŸ„ Healthcare

⚕ Healthcare Insurance

💰 $2.8M Seed Round on 2022-09

Healthcare ‱ Healthcare Insurance ‱ Mental Health

MEDvidi is an online mental health treatment center aiming to make professional care accessible and affordable for everyone. The healthcare experts at MEDvidi provide personalized treatment plans for a variety of mental health conditions, including ADHD, anxiety, depression, insomnia, and OCD. With services that include initial assessments, ongoing support, and medication management through virtual consultations, MEDvidi seeks to enhance mental wellness with compassionate care tailored to individual needs.

📋 Description

‱ Set and drive the technical vision and quarterly roadmap for infrastructure, with clear trade-offs and measurable goals ‱ Run and evolve AWS and Kubernetes (EKS) infrastructure, including cluster management, autoscaling, policy enforcement, and zero-downtime operations ‱ Own Infrastructure as Code using Terraform and AWS CDK in TypeScript ‱ Own GitLab CI/CD, including reusable/shared templates, OIDC, and self-managed GitLab ‱ Build and own practical observability using Prometheus, Grafana, OpenTelemetry, OpenSearch, and CloudWatch ‱ Maintain log pipelines and APM as a shared view of system health ‱ Operate blue-green deployments with health-gated automated rollback ‱ Own zero-downtime PostgreSQL schema migrations and CI migration gating ‱ Lead security engineering in a HIPAA environment, including secrets hygiene, PHI-aware logging and data handling, leak scanning, and Vault managed as code ‱ Partner directly with product teams to remove infrastructure friction and improve developer experience ‱ Use agentic AI as a core workflow, integrating autonomous-agent output into production ‱ Set technical direction, ship infrastructure changes hands-on, and own reliability, cost, security, performance, deployment health, and developer velocity outcomes

🎯 Requirements

‱ 6+ years of experience in DevOps/infrastructure engineering ‱ Strong systems fundamentals, Linux administration, and troubleshooting, including performance analysis, resource management, and process debugging ‱ Hands-on AWS experience with EC2, EKS, RDS, ElastiCache, Lambda, SQS, EventBridge, API Gateway, ALB, and S3 ‱ Production Kubernetes/EKS experience, including cluster management, node scaling, and policy enforcement; Karpenter, Kyverno, or similar ‱ Strong Infrastructure as Code experience with Terraform and AWS CDK in TypeScript ‱ CI/CD ownership with GitLab CI/CD, reusable/shared templates, OIDC id_tokens, and self-managed GitLab ‱ Practical monitoring and observability experience with Prometheus, Grafana, OpenTelemetry, OpenSearch, and CloudWatch; log shipping and error tracking/APM ‱ Practical security engineering experience with secrets rotation, short-lived credentials, leak scanning, and PHI-aware logging ‱ HashiCorp Vault as code experience, including KV, JWT/OIDC authentication for CI, and policy design ‱ Experience with blue-green deployments and automated, health-gated rollback ‱ PostgreSQL zero-downtime schema migration experience using expand/contract and migration gating in CI ‱ Containers experience with Docker, ECR, immutable tags, and image lifecycle ‱ Network and protocol fundamentals, including load balancing, TLS, and DNS ‱ Hands-on agentic AI workflows, such as Claude Code or similar, including delegating to autonomous agents and integrating their output into production ‱ Developer-focused mindset and strong problem-solving for complex system issues ‱ Strong technical writing, including docs-as-code, ADRs, and design documents via MRs ‱ Fluent Russian and English (B2) ‱ Experience working effectively in remote, distributed teams ‱ Preferred: experience in HIPAA, SOC 2, or similar regulated environments; Ansible; Node.js operations; GitOps tooling; deeper PostgreSQL administration; AWS certifications

đŸ–ïž Benefits

‱ Competitive compensation package ‱ Fully remote long-term collaboration under a B2B model ‱ Health insurance after the probation period ‱ Sports & wellness compensation ‱ Personalized English lessons via Preply ‱ 19 paid vacation days annually ‱ 4 additional wellness days each year ‱ Paid sick leave for the first 5 working days ‱ Thoughtful gifts for key life events ‱ Offline corporate events

Apply Now

Similar Jobs

🕒 July 28

Wand AI

51 - 200

đŸ€– Artificial Intelligence

🏱 Enterprise

☁ SaaS

Senior Staff SRE Engineer designing and operating scalable infrastructure for AI-driven products. Collaborating across teams to improve reliability and performance with a focus on SRE practices.

AWS

Azure

Cloud

Distributed Systems

Kubernetes

Terraform

🕒 July 28

Wand AI

51 - 200

đŸ€– Artificial Intelligence

🏱 Enterprise

☁ SaaS

Head of SRE establishing and scaling Site Reliability Engineering function at Wand. Driving operational excellence and improving production stability for AI products.

AWS

Azure

Cloud

Kubernetes

Terraform

🕒 May 20

Replit

51 - 200

đŸ€– Artificial Intelligence

đŸ€ B2B

Join Replit as a Staff Site Reliability Engineer, enhancing performance and reliability of our infrastructure. Collaborate to ensure scalable solutions while mentoring engineers.

Cloud

Distributed Systems

Kubernetes

Python

Terraform

Go

🕒 February 17

Thrill

11 - 50

🎼 Gaming

đŸ„œ AR/VR

Infrastructure/DevOps Engineer responsible for managing AWS and Kubernetes at Thrill Labs. Working on high-scalability projects and improving security measures in a fast-growing tech startup.

AWS

Kubernetes

Linux

Postgres

Python