Python Inference Engineer

🔥 13 hours ago

🌐 Cyprus, Poland, +1 more countries – Remote

infoinfo

⏰ Full Time

🟡 Mid-level

🟠 Senior

🔙 Backend Engineer

👻 Ghost score 10%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Gcore

Gcore

201 - 500 employees

💼 Consulting

📦 Logistics

📣 Marketing

Consulting • Logistics • Marketing

Gcore is a global provider of cloud, edge, and AI solutions that accelerate AI training, deliver comprehensive cloud services, enhance content delivery, and protect servers and applications. With over 180 points of presence worldwide and a network capacity of 200+ Tbps, Gcore offers secure, flexible, and scalable infrastructure services. Its integrated offerings, including Edge Cloud, Edge Network, Edge Security, and AI Infrastructure, are designed to meet the needs of businesses looking to scale and control their global infrastructure efficiently. Gcore also provides robust DDoS protection and origin shielding to ensure uninterrupted online operations, making it a trusted partner for thousands of businesses worldwide.

📋 Description

• Build and improve the inference layer of the Gcore Inference platform • Integrate and operate inference frameworks such as vLLM, SGLang, NVIDIA Dynamo, and TensorRT-LLM • Bring new language and multimodal models into production • Improve inference latency, throughput, memory use, GPU utilization, and cost efficiency • Debug performance and reliability issues across model code, inference frameworks, GPU execution, networking, and Kubernetes • Work with platform, infrastructure, product, and customer-facing teams to turn inference improvements into reliable product features • Contribute improvements to open-source inference projects when appropriate

🎯 Requirements

• 5+ years of experience writing reliable, well-tested production code • Strong Python skills and experience designing production systems • Hands-on experience with PyTorch and deploying machine learning models • Experience with Linux, Docker, and Kubernetes • Experience in at least one relevant area: distributed systems, GPU computing, ML runtimes, model optimization, or cluster scheduling • Ability to debug complex problems across software, infrastructure, and hardware • Strong sense of developer experience • Genuine interest in inference engineering and motivation to learn • Good communication and collaboration skills • Nice to have: experience with vLLM, SGLang, NVIDIA Dynamo, TensorRT-LLM, or a similar inference framework • Nice to have: experience running GPU workloads in production • Nice to have: knowledge of quantization, continuous batching, speculative decoding, prefix caching, chunked prefill, or LoRA serving • Nice to have: experience with CUDA, Triton, TensorRT, or other GPU programming tools • Nice to have: experience profiling and improving model latency, throughput, memory use, or GPU utilization • Nice to have: experience with distributed inference, multi-GPU systems, scheduling, or autoscaling • Nice to have: contributions to open-source ML, inference, or infrastructure projects

🏖️ Benefits

• Competitive compensation • Flexible working hours and hybrid or remote options, depending on your role • Work from anywhere in the world for up to 45 days per year • Private medical insurance for you and your family* • Extra paid vacation and sick leave days* • Support for life’s important moments and celebrations • Language courses to help you connect and grow • Modern, welcoming offices with snacks, drinks, and entertainment* • Team sports and social activities*

Apply Now

Similar Jobs

🕒 August 21

Aiphoria

51 - 200

💼 Consulting

📦 Logistics

📣 Marketing

Backend Team Lead building scalable LLM platform services for award-winning AI products. Leading backend execution, agent framework development, integrations, and platform reliability.

Distributed Systems

Grafana

GRPC

Kafka

Kubernetes

Postgres

Python

🕒 August 18

Fundraise Up

51 - 200

🤲 Charity

💳 Fintech

☁️ SaaS

Senior Backend Engineer scaling Fundraise Up’s global fundraising platform for nonprofits. Building Node.js and TypeScript services across checkout, donor portals, and analytics systems.

🗣️🇷🇺 Russian Required

ElasticSearch

JavaScript

Kafka

MongoDB

Node.js

NoSQL

RabbitMQ

Redis

TypeScript

🕒 June 22

Holiston Media Ltd

1 - 10

🤝 B2B

💸 Finance

🛍️ eCommerce

Python Developer at primexbt developing CFD trading tools and integrations with liquidity providers. Responsible for optimizing trading workflows and maintaining connectivity with external partners.

Python

🕒 April 23

Mayflower

501 - 1000

🥽 AR/VR

🤖 Artificial Intelligence

📱 Media

Kotlin/Java Developer focusing on developing and optimizing payment integrations. Collaborating with teams to enhance existing services and ensuring high-quality code delivery.

Java

Kotlin

MySQL

NoSQL

Redis

Spring

🕒 April 22

Mayflower

501 - 1000

🥽 AR/VR

🤖 Artificial Intelligence

📱 Media

Kotlin/Java Developer involved in developing payment gateways systems with multiple integrations and high-security compliance. Responsible for designing, implementing, and optimizing payment services and APIs.

Apache

Kafka

Kotlin

MySQL

NoSQL

React

Redis

Spring

Svelte