LLM Inference Engineer

πŸ”₯ 0 minutes ago

Apply Now
Find Similar Remote Jobs

πŸ“Š Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of NEAR AI

NEAR AI

11 - 50 employees

Founded 2017

πŸ€– Artificial Intelligence

πŸ”’ Cybersecurity

🏒 Enterprise

Artificial Intelligence β€’ Cybersecurity β€’ Enterprise

NEAR AI is a privacy-first AI infrastructure company that runs confidential inference and autonomous AI agents for enterprises, governments, and AI applications. Its NEAR AI Cloud and IronClaw agent framework execute models and workflows inside hardware-enforced Trusted Execution Environments (TEEs) with cryptographic isolation, zero operator access, and hardware-signed attestation for independent verification. NEAR AI offers a turnkey platform for deploying private models, automating repetitive operations, integrating with internal tools (Slack, email, Notion), and serving regulated or sovereign workloads via a marketplace of agents for finance, legal, ops, research and security.

πŸ“‹ Description

β€’ Architect and maintain production high-traffic LLM serving systems. β€’ Optimize throughput, latency, and cost for leading open-source LLMs.

🎯 Requirements

β€’ Strong hands-on experience in LLM inference, with expertise debugging and optimizing major inference engines such as SGLang, vLLM, or TensorRT. β€’ Deep knowledge of state-of-the-art GPU architectures, and effectively exploit them using PyTorch, Triton, CuTe, CUDA, etc. β€’ Proven track record in designing and maintaining end-to-end high-traffic LLM serving systems. β€’ Strong problem-solving skills and ability to communicate technical ideas clearly.

Apply Now

Similar Jobs

πŸ”₯ 6 hours ago

Century Interactive

51 - 200

πŸ“£ Marketing

Senior Machine Learning Engineer specializing in NLP and LLM-powered models at Call Box. Lead the design and deployment of scalable AI-driven solutions in an exciting collaborative environment.

AWS

Azure

Cloud

Docker

Google Cloud Platform

Kubernetes

Python

PyTorch

Tensorflow

πŸ”₯ 23 hours ago

InnoData

2 - 10

🀝 B2B

πŸ’Ό Consulting

🌍 Social Impact

AI/ML Research Engineer designing and implementing LLM training and evaluation pipelines. Collaborating with technical teams to improve foundation model performance at Innodata.

πŸ‡ΊπŸ‡Έ United States – Remote

πŸ’΅ $80k - $175k / year

⏰ Full Time

🟒 Junior

🟑 Mid-level

πŸ—£οΈ LLM Engineer

Python

PyTorch

Tensorflow

πŸ•’ Yesterday

EWOR

201 - 500

πŸ“š Education

πŸ’Έ Finance

πŸ’Ό Consulting

Co-Founder / CEO responsible for launching and scaling AI Infrastructure startups in a supportive environment. Engage with experienced entrepreneurs while developing a high-growth venture.

πŸ•’ Yesterday

EWOR

201 - 500

πŸ“š Education

πŸ’Έ Finance

πŸ’Ό Consulting

Co-Founder / CCO leading an AI Infrastructure startup backed by experienced entrepreneurs. Building startup with coaching and funding support to reach significant revenues.

πŸ•’ 5 days ago

Honeycomb.io

51 - 200

☁️ SaaS

🏒 Enterprise

πŸ€– Artificial Intelligence

Senior Software Engineer II at Honeycomb.io building observability solutions for AI workloads with tech stack in Go and React. Collaborating across teams to enhance product features with clear communication.

React

TypeScript