DevOps Engineer, AI Inference

🔥 0 minutes ago

🌐 Cyprus, Poland, +3 more countries – Remote

infoinfo

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

👻 Ghost score 10%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Gcore

Gcore

201 - 500 employees

💼 Consulting

📦 Logistics

📣 Marketing

Consulting • Logistics • Marketing

Gcore is a global provider of cloud, edge, and AI solutions that accelerate AI training, deliver comprehensive cloud services, enhance content delivery, and protect servers and applications. With over 180 points of presence worldwide and a network capacity of 200+ Tbps, Gcore offers secure, flexible, and scalable infrastructure services. Its integrated offerings, including Edge Cloud, Edge Network, Edge Security, and AI Infrastructure, are designed to meet the needs of businesses looking to scale and control their global infrastructure efficiently. Gcore also provides robust DDoS protection and origin shielding to ensure uninterrupted online operations, making it a trusted partner for thousands of businesses worldwide.

📋 Description

• Design, develop, and maintain infrastructure for AI inference workloads, including GPU scheduling, model deployment pipelines, and data access patterns in on-prem environments • Build and manage monitoring and observability tools for AI inference platforms, including dashboards, alerts, and runbooks for model health and system performance • Collaborate with ML engineers and platform teams to design system architecture for AI workloads • Integrate inference runtimes • Test AI inference performance at scale

🎯 Requirements

• This position is available only under an employment (labor) agreement • Strong understanding of Kubernetes architecture, including CNI, CSI, operators, ingress/gateway, and control plane components • Hands-on experience operating and troubleshooting production Kubernetes clusters • Strong Linux and networking troubleshooting skills, including DNS, routing, firewalling, TLS, MTU, connectivity and performance issues • Ability to develop automation and operational tooling using Python, Go, or Bash • Experience with Terraform, Ansible, or similar IaC/configuration management tools • Experience with VictoriaMetrics/Grafana or similar monitoring, alerting, and troubleshooting tools • Strong experience with Git-based workflows and CI/CD pipelines • Familiarity with Cluster API or similar Kubernetes cluster lifecycle management technologies • Hands-on operation or administration of Slurm clusters • Knowledge of Argo CD, GitOps workflows, Helm, or Helmfile • Background working with managed platforms, PaaS, or cloud services • Exposure to bare metal, GPU, HPC, or other high-performance computing environments • Familiarity with the NVIDIA GPU stack, RDMA/InfiniBand, or high-performance networking • Knowledge of OpenStack or similar cloud infrastructure platforms • Hands-on experience developing Kubernetes operators or controllers

🏖️ Benefits

• Competitive compensation • Flexible working hours and hybrid or remote options, depending on your role • Work from anywhere in the world for up to 45 days per year • Private medical insurance for you and your family* • Extra paid vacation and sick leave days* • Support for life’s important moments and celebrations • Language courses to help you connect and grow • Modern, welcoming offices with snacks, drinks, and entertainment* • Team sports and social activities*

Apply Now

Similar Jobs

🕒 Yesterday

Vivid Money

201 - 500

💳 Fintech

🏦 Banking

💸 Finance

DevSecOps Engineer securing AWS, Kubernetes, and AI infrastructure at Vivid, a fintech platform serving businesses and individuals across Europe. Improving cloud, application, and vulnerability security.

AWS

Cloud

Kubernetes

SDLC

🕒 Yesterday

Vivid Money

201 - 500

💳 Fintech

🏦 Banking

💸 Finance

DevSecOps Engineer securing AWS, Kubernetes, and AI infrastructure for Vivid, a European fintech platform. Improving cloud security, vulnerability management, and AI-aware controls across engineering systems.

AWS

Cloud

Kubernetes

SDLC

🕒 July 21

Fundraise Up

51 - 200

🤲 Charity

💳 Fintech

☁️ SaaS

DevOps Engineer managing observability and CI/CD for Fundraise Up, a global fundraising platform. Driving technical initiatives and mentoring less experienced engineers in a complex product ecosystem.

🗣️🇷🇺 Russian Required

Ansible

Docker

Grafana

Jenkins

Kubernetes

Linux

Prometheus

Python

🕒 June 24

Aylo

1001 - 5000

👥 B2C

🎮 Gaming

Sr DevOps in high traffic environment supporting production and development infrastructure for diverse teams. Collaborating on architecture improvements and optimizing workflows with modern tools.

Ansible

Cloud

DNS

Kubernetes

Linux

Python

VMware