Senior Inference Engineer

Job not on LinkedIn

🔥 0 minutes ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of vCluster

vCluster

51 - 200 employees

☁️ SaaS

🏢 Enterprise

💰 $24M Series A on 2024-05

SaaS • Enterprise

vCluster is a Kubernetes virtualization product (from loft-sh) that creates virtual Kubernetes clusters on top of existing Kubernetes infrastructure. It enables dedicated, multi-tenant or private-node clusters for use cases like internal platform standardization, hybrid Kubernetes deployments, sovereign or dedicated customer environments, AI cloud providers and distributed inference. vCluster is positioned as a tooling/product solution for enterprises operating Kubernetes at scale, offering managed features, releases and integrations for internal K8s platforms and cloud-native workflows.

📋 Description

• Deploy LLMs to production across one or more machines on GPU infrastructure • Own the full pipeline from a customer query to the served response • Stand up and operate serving infrastructure using vLLM, SGLang, or TensorRT-LLM • Optimize inference at scale using quantization, batching, caching, and routing to manage latency and cost • Build production infrastructure in Python or Golang • Build the first inference-platform iteration alongside the CTO • Lead the inference platform roadmap and partner with Product on future development • Explain technical concepts clearly to engineers and non-technical stakeholders

🎯 Requirements

• Production experience deploying and serving LLMs using vLLM, SGLang, or TensorRT-LLM • Hands-on experience with quantization, batching, caching, and routing • Strong production engineering skills in Python or Golang • Strong communication skills, including explaining technical concepts to technical and non-technical stakeholders • Familiarity with Docker and Kubernetes • Hands-on generative AI experience with PyTorch and Transformers • Understanding of the GPU stack, including CUDA, NCCL, drivers, and related libraries • Knowledge of model architectures and fine-tuning approaches • Experience with NVIDIA Dynamo

🏖️ Benefits

• Competitive compensation package, including equity • Health, dental, vision, and life insurance • Insurance plans for you and eligible dependents (benefits vary depending on country) • Flexible working schedule • Workplace flexibility • Remote-first work culture • Bonus compensation is offered

Apply Now

Similar Jobs

🔥 29 minutes ago

Hewlett Packard Enterprise

10,000+ employees

🏢 Enterprise

🔧 Hardware

☁️ SaaS

HPE Resident Engineer supporting Meta’s Juniper/HPE service-provider network. Designing, testing, troubleshooting, and implementing advanced routing and network solutions remotely.

🔥 38 minutes ago

Hewlett Packard Enterprise

10,000+ employees

🏢 Enterprise

🔧 Hardware

☁️ SaaS

HPE Resident Engineer supporting Meta’s Juniper/HPE service-provider networks. Designing, testing, troubleshooting, and upgrading customer network infrastructure.

🔥 49 minutes ago

webAI™

51 - 200

🏭 Manufacturing

📦 Logistics

🏥 Healthcare

Embedded signal processing engineer deploying secure, distributed AI systems on rugged government edge hardware at webAI. Designing multimodal fusion, networking, simulation, calibration, and field-testing solutions for disconnected environments.

🔥 53 minutes ago

1Password

501 - 1000

🔒 Cybersecurity

☁️ SaaS

⚡ Productivity

Privacy Engineer building AI-assisted controls and workflows for 1Password, a cybersecurity SaaS company. Improving privacy reviews, data handling, telemetry, retention, and deletion at scale.

🔥 1 hour ago

Leidos

10,000+ employees

🏥 Healthcare

💼 Consulting

📦 Logistics

Lead Microgrid Engineer designing solar, storage, generation, and controls projects for Leidos’ utility and commercial customers. Providing Owner’s Engineering, interconnection, commissioning, and technical review leadership.

Flash