Member of Technical Staff, Inference

đŸ”„ 0 minutes ago

🌏 Anywhere in the World

⏰ Full Time

🔮 Lead

đŸ–„ Software Engineer

đŸ‘» Ghost score 13%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Inferact

Inferact

11 - 50 employees

Founded 2025

đŸ€– Artificial Intelligence

đŸ€ B2B

🏱 Enterprise

Artificial Intelligence ‱ B2B ‱ Enterprise

Inferact is a startup founded by the creators and core maintainers of vLLM, the leading open-source LLM inference engine. The company aims to accelerate AI progress by making model inference cheaper and faster, expanding vLLM's performance and support for emerging architectures and accelerator hardware. Inferact combines deep expertise at the intersection of models and hardware to provide inference infrastructure used by research labs, hyperscalers, and startups, while continuing to develop and contribute optimizations back to the open-source community.

📋 Description

‱ Push the boundaries of LLM and diffusion model serving ‱ Work at the core of vLLM to optimize model execution across diverse hardware and architectures ‱ Develop inference runtime innovations for mixture-of-experts, multimodal, and agentic architectures ‱ Implement inference techniques and model architectures from research papers ‱ Contribute performant and maintainable code ‱ Debug complex machine-learning codebases ‱ Directly improve how AI inference is run by making inference cheaper and faster

🎯 Requirements

‱ Bachelor's degree or equivalent experience in computer science, engineering, or similar ‱ Deep understanding of transformer architectures and their variants ‱ Strong programming skills in Python with experience in PyTorch internals ‱ Experience with LLM inference systems such as vLLM, TensorRT-LLM, SGLang, or TGI ‱ Ability to read and implement model architectures and inference techniques from research papers ‱ Ability to contribute performant and maintainable code and debug in complex ML codebases ‱ Preferred: deep understanding of KV-cache memory management, prefix caching, and hybrid model serving ‱ Preferred: familiarity with RL frameworks and algorithms for LLMs ‱ Preferred: experience with multimodal inference across audio, image, video, and text ‱ Contributions to open-source ML or system infrastructure projects are preferred ‱ Bonus: core feature implementation in vLLM or other inference engine projects ‱ Bonus: contributions to vLLM integrations such as verl, OpenRLHF, Unsloth, or LlamaFactory ‱ Bonus: widely-shared technical blogs or side projects on vLLM or LLM inference ‱ Required application materials: resume, GitHub handle, and a link to a relevant personal project, open-source contribution, or technical blog post

đŸ–ïž Benefits

‱ Competitive benefits appropriate to your location, including health coverage where applicable ‱ Equity ‱ Visa sponsorship on a case-by-case basis ‱ Fully remote work ‱ Timezone-flexible schedule with regular overlap with Pacific Time for critical syncs

Apply Now

Similar Jobs

🕒 August 19

Progress Partners

51 - 200

☁ SaaS

⚡ Productivity

đŸ€– Artificial Intelligence

Engineering Delivery Manager improving predictability, quality, and visibility across external engineering teams. Progress Partners builds SaaS and mobile platforms used worldwide.

Azure

🕒 August 2

Supabase

51 - 200

☁ SaaS

🔌 API

đŸ€– Artificial Intelligence

Engineering Productivity Engineer at Supabase enhancing local development workflows and optimizing engineering processes. Focus on developer productivity with AI-assisted tools and CI systems.

Python

Rust

TypeScript

Go

🕒 May 19

Supabase

51 - 200

☁ SaaS

🔌 API

đŸ€– Artificial Intelligence

Engineer driving the evolution of OrioleDB and collaborating with the PostgreSQL community at Supabase. Design and implement new database features and ensure system reliability.

Open Source

Postgres

🕒 March 31

ClosedWon

1 - 10

đŸ’Œ Consulting

📩 Logistics

📣 Marketing

VP of Engineering leading the technological vision and strategy at ClosedWon, an AI sales coaching company. Collaborating with leadership to drive and implement innovative solutions.

AWS

JavaScript

React

Redux

Ruby

Ruby on Rails