Inference Infrastructure Architect

🔥 20 hours ago

🇨🇳 China – Remote

⏰ Full Time

🟡 Mid-level

🟠 Senior

👷 Infrastructure Engineer

👻 Ghost score 11%

infoinfo

🗣️🇨🇳 Chinese Required

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Telnyx

Telnyx

201 - 500 employees

Founded 2015

💼 Consulting

📦 Logistics

📡 Telecommunications

💰 $2.1M Seed Round on 2014-08

Consulting • Logistics • Telecommunications

Telnyx is a modular, cloud-native platform that enables the creation of custom communications and connectivity applications. It offers a comprehensive suite of APIs accessible through intuitive SDKs and web tools. Telnyx provides AI-powered connectivity solutions, programmable networking, IoT SIM cards, and private global networks for high-quality communication services. Their platform is designed to improve customer experience through AI-powered voice solutions and offers low-latency AI services using GPU-powered infrastructure. Telnyx supports various industries by providing tools for building secure, scalable, and innovative communication and connectivity solutions worldwide.

📋 Description

• Operate and expand Telnyx’s B300 GPU fleet efficiently • Maximize inference throughput per GPU-dollar while meeting latency and reliability SLOs • Continuously reduce cost per token • Build serverless serving pools using vLLM/SGLang, continuous batching, prefix caching, low-precision serving, and MoE expert parallelism • Build the fleet layer with llm-d/NVIDIA Dynamo on Kubernetes Gateway API, KV-cache-aware routing, prefill/decode disaggregation, KV tiering, and multi-LoRA serving • Build Kubernetes-on-bare-metal infrastructure with GPU Operator, topology-aware scheduling, LeaderWorkerSet, Kueue, Kata Containers, and OpenStack Ironic • Implement model distribution, warm pools, and inference-metric-driven autoscaling • Develop dedicated tenant pools, GPU-hour metering, latency SLOs, private networking, adapter versioning, canary rollouts, and rollback processes • Integrate observability using DCGM, Prometheus, and OpenTelemetry • Perform capacity planning using roofline analysis, batching curves, utilization, and cost-per-token metrics • Select, benchmark, and integrate production components • Contribute upstream to the open-source serving stack • Provide a brief application example describing an improved inference system, bottleneck, intervention, and measured result

🎯 Requirements

• Production LLM serving experience under meaningful traffic and latency constraints • Experience operating Kubernetes on GPU fleets end to end, including GPU Operator, device plugins, node pools, topology-aware placement, gang scheduling, GitOps rollouts, and pod-placement debugging • Deep operational command of vLLM or SGLang, including deployment, tuning, upgrades, TP/EP parallelism, quantization, batching, KV-cache settings, prefix caching, and disaggregation • System-level performance engineering using engine and DCGM metrics, roofline reasoning, and numerical deployment sizing • Python and Go for automation • Strong Linux, networking, and storage knowledge • Ability to identify CUDA-related problems and route them upstream • Ability to work with the community in English and Chinese • Ability to write effective runbooks and design documents • Experience with inference platforms at large-scale organizations is valued • Contribution to or heavy production use of vLLM, SGLang, llm-d, NVIDIA Dynamo, Ray/KubeRay, Kata Containers, Volcano, Kueue, HAMi, Dragonfly, Mooncake, or Envoy is valued • CNCF/OpenInfra community presence is valued • Fine-tuning/RL infrastructure, real-time voice latency work, and multi-region/data-residency experience are nice to have

🏖️ Benefits

• A globally expanding B300 fleet operated end to end • Greenfield inference platform ownership and opportunity to help build the China team • Open-source-first environment • Conference travel support • Remote work with a global, async-friendly team • No relocation required • Visa sponsorship available if relocating to hiring entities in the Netherlands, United States, Ireland, or Saudi Arabia

Apply Now

Similar Jobs

🕒 June 29

NVIDIA

10,000+ employees

🏥 Healthcare

🏭 Manufacturing

🤖 Artificial Intelligence

Engineer for CI/CD infrastructure, enhancing reliability and performance for NVIDIA's deep learning compilers. Collaborating with teams to automate workflows across GPU and accelerator environments.

🇨🇳 China – Remote

⏰ Full Time

🟡 Mid-level

🟠 Senior

👷 Infrastructure Engineer

Distributed Systems

Jenkins

Python