Software Engineer, Golang, Slurm

🔥 0 minutes ago

🌐 Cyprus, Poland, +2 more countries – Remote

infoinfo

⏰ Full Time

🟡 Mid-level

🟠 Senior

🧑‍💻 Full-stack Engineer

👻 Ghost score 10%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Gcore

Gcore

201 - 500 employees

💼 Consulting

📦 Logistics

📣 Marketing

Consulting • Logistics • Marketing

Gcore is a global provider of cloud, edge, and AI solutions that accelerate AI training, deliver comprehensive cloud services, enhance content delivery, and protect servers and applications. With over 180 points of presence worldwide and a network capacity of 200+ Tbps, Gcore offers secure, flexible, and scalable infrastructure services. Its integrated offerings, including Edge Cloud, Edge Network, Edge Security, and AI Infrastructure, are designed to meet the needs of businesses looking to scale and control their global infrastructure efficiently. Gcore also provides robust DDoS protection and origin shielding to ensure uninterrupted online operations, making it a trusted partner for thousands of businesses worldwide.

📋 Description

• Design and build a managed Slurm service on Kubernetes • Write clean, reliable, and maintainable Go code • Develop scheduling and orchestration capabilities for GPU-intensive and distributed workloads • Build observability and automated remediation for GPU, node, network, and control-plane failures using VictoriaMetrics, Grafana, and DCGM • Preserve traditional Slurm cluster behavior while running infrastructure on Kubernetes • Diagnose performance and reliability issues across GPUs, schedulers, hardware, high-performance networks, and distributed storage systems • Take end-to-end ownership of complex distributed-system challenges • Collaborate with technology partners and a global team building infrastructure and software for AI, cloud, network, and security

🎯 Requirements

• Hands-on experience using Slurm in production from a user’s perspective, including submitting and debugging workloads with sbatch, srun, squeue, and sinfo • Strong proficiency in Go • Experience building production-grade Kubernetes operators, controllers, CRDs, and reconciliation loops • Experience preserving traditional Slurm cluster behavior while running the underlying infrastructure on Kubernetes • Experience diagnosing performance and reliability issues across GPUs, schedulers, hardware, high-performance networks, and distributed storage systems • Product mindset and strong customer empathy • Excellent communication skills and ability to take end-to-end ownership of complex distributed-system challenges • Nice to have: experience operating large-scale HPC or GPU clusters for external customers • Nice to have: experience with PyTorch distributed training and other large-scale AI/ML frameworks • Nice to have: experience with InfiniBand, RoCE, RDMA, GPUDirect, Lustre, WEKA, Ceph, or similar high-performance infrastructure • Nice to have: experience building unified job-submission workflows across Kubernetes and Slurm • Nice to have: experience in GPU-cloud or HPC product engineering environments • Nice to have: contributions to Slurm, Kubernetes, Soperator, or other cloud-native and HPC open-source projects

🏖️ Benefits

• Competitive compensation • Flexible working hours • Hybrid or remote options, depending on your role • Work from anywhere in the world for up to 45 days per year • Private medical insurance for you and your family* • Extra paid vacation and sick leave days* • Support for life’s important moments and celebrations • Language courses • Modern, welcoming offices with snacks, drinks, and entertainment* • Team sports and social activities*

Apply Now

Similar Jobs

🕒 2 days ago

Miratech

501 - 1000

🤝 B2B

💼 Consulting

☁️ SaaS

Software Architect defining scalable cloud-native architectures for Miratech, a global IT services and consulting company. Leading technology strategy, PoCs, and modernization across enterprise platforms.

AWS

Azure

Cloud

Distributed Systems

Google Cloud Platform

Kubernetes

Microservices

🕒 August 10

TalentJar

1 - 10

🎯 Recruiter

👥 HR Tech

🤝 B2B

Full-stack engineer building reliable payment orchestration applications with TypeScript, React, Node.js, and PostgreSQL. Owning features end to end while shaping safe AI-assisted engineering practices.

JavaScript

Next.js

Node.js

Postgres

React

TypeScript

🕒 August 3

Fundraise Up

51 - 200

🤲 Charity

💳 Fintech

☁️ SaaS

Senior Fullstack Developer creating scalable donation widgets and portals for Fundraise Up. Join a global team focused on high-impact engineering and innovation.

🗣️🇷🇺 Russian Required

JavaScript

Node.js

React

Vue.js

Webpack

🕒 June 23

GoMining

201 - 500

💼 Consulting

📣 Marketing

₿ Crypto

Lead FullStack Engineer with strong backend expertise to help build high-quality products while leading a small engineering team. Leverage AI tools for engineering quality and development velocity.

Angular

Distributed Systems

Docker

JavaScript

Kubernetes

Node.js

Postgres

RabbitMQ

React

Redis

TypeScript

Web3

🕒 June 20

GoMining

201 - 500

💼 Consulting

📣 Marketing

₿ Crypto

FullStack Engineer focused on back-end development, building enterprise products with Product and Engineering teams. Responsibilities include front-end applications and backend service implementation.

Angular

Distributed Systems

Docker

JavaScript

Jest

Kubernetes

Linux

Next.js

Node.js

Postgres

React

Redis

SCSS

TypeScript

Web3

Webpack