AI Infrastructure Engineer

Job not on LinkedIn

🔥 3 minutes ago

🇸🇬 Singapore – Remote

💵 S$155k - S$175k / year

⏰ Full Time

🟡 Mid-level

🟠 Senior

👷 Infrastructure Engineer

👻 Ghost score 3%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of vCluster

vCluster

51 - 200 employees

☁️ SaaS

🏢 Enterprise

💰 $24M Series A on 2024-05

SaaS • Enterprise

vCluster is a Kubernetes virtualization product (from loft-sh) that creates virtual Kubernetes clusters on top of existing Kubernetes infrastructure. It enables dedicated, multi-tenant or private-node clusters for use cases like internal platform standardization, hybrid Kubernetes deployments, sovereign or dedicated customer environments, AI cloud providers and distributed inference. vCluster is positioned as a tooling/product solution for enterprises operating Kubernetes at scale, offering managed features, releases and integrations for internal K8s platforms and cloud-native workflows.

📋 Description

• Lead end-to-end technical deployments for GPU neocloud and AI Factory customers, from bare metal configuration to validated vCluster environments • Configure and troubleshoot bare metal GPU node infrastructure, including CNI, GPU Operator, distributed storage backends, and RDMA/InfiniBand • Deploy and validate Kubernetes and vCluster for GPU-powered managed Kubernetes • Work alongside customer teams to build operational self-sufficiency • Document reusable playbooks and deployment architectures • Collaborate with Engineering and Product to surface infrastructure challenges and inform the roadmap • Join Sales in pre-sales proof-of-value engagements requiring deep infrastructure expertise

🎯 Requirements

• 5+ years of experience deploying and operating Kubernetes in production • Practical knowledge of NVIDIA GPU Operators, CUDA tooling, and systems-level configuration for GPU nodes • Deep understanding of CNI plugins, overlay networks, load balancing, and connectivity diagnosis in layered environments • Experience with persistent volume configuration, CSI drivers, and distributed systems such as Ceph, Rook, Weka, or Longhorn • Comfort operating in ambiguous, fast-moving environments • Experience with Bash, Python, or Go automation scripts is a bonus • CKA certification or experience writing Kubernetes Operators is a bonus • Experience with inference serving, GPU scheduling, and LLM deployment tooling is a bonus • Experience building AI Automation in documentation is a bonus

🏖️ Benefits

• Equity • Bonus • Health insurance • Dental insurance • Vision insurance • Life insurance • Plans for eligible dependents • Flexible working schedule • Workplace flexibility • Remote-first culture

Apply Now

Similar Jobs

🕒 August 11

Cohere

11 - 50

🤖 Artificial Intelligence

🏢 Enterprise

☁️ SaaS

Forward Deployed Engineer deploying Cohere’s North enterprise AI workspace for customer infrastructure. Integrating secure AI solutions across private cloud and on-premises environments in Singapore.

🗣️🇰🇷 Korean Required

AWS

Azure

Cloud

Google Cloud Platform

Kubernetes

🕒 July 27

GetBlock

11 - 50

₿ Crypto

🔌 API

🌐 Web 3

Infrastructure Engineer designing and maintaining blockchain systems for global Web3 and AI infrastructure. Resolving API performance issues and ensuring reliability with a focus on stability and security.

Grafana

GraphQL

GRPC

Linux

Node.js

Python

🕒 May 28

Pioneer

201 - 500

🛡️ Insurance

💼 Consulting

🏦 Banking

Infrastructure Development Engineer focused on building and optimizing the BNB Chain Middleware Stack. Simplifying Web3 development for decentralized applications with a tech team.

JavaScript

Next.js

React

TypeScript

Vue.js

Web3