LLM Inference Engineer

🕒 July 28

🏄 California – Remote

infoinfo

⏰ Full Time

🟡 Mid-level

🟠 Senior

👷🏻‍♀️ Engineer

👻 Ghost score 30%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of NEAR AI

NEAR AI

11 - 50 employees

Founded 2017

🤖 Artificial Intelligence

🔒 Cybersecurity

🏢 Enterprise

Artificial Intelligence • Cybersecurity • Enterprise

NEAR AI is a privacy-first AI infrastructure company that runs confidential inference and autonomous AI agents for enterprises, governments, and AI applications. Its NEAR AI Cloud and IronClaw agent framework execute models and workflows inside hardware-enforced Trusted Execution Environments (TEEs) with cryptographic isolation, zero operator access, and hardware-signed attestation for independent verification. NEAR AI offers a turnkey platform for deploying private models, automating repetitive operations, integrating with internal tools (Slack, email, Notion), and serving regulated or sovereign workloads via a marketplace of agents for finance, legal, ops, research and security.

📋 Description

• Architect and maintain production high-traffic LLM serving systems • Optimize throughput, latency, and cost for leading open-source LLMs • Push the boundaries of how large language models are served • Contribute to decentralized and confidential machine learning infrastructure enabling user-owned AI • Help build highly scalable and efficient infrastructure for open-source AI at global scale

🎯 Requirements

• Strong hands-on experience in LLM inference, with expertise debugging and optimizing major inference engines such as SGLang, vLLM, or TensorRT • Deep knowledge of state-of-the-art GPU architectures • Ability to effectively exploit GPU architectures using PyTorch, Triton, CuTe, CUDA, etc. • Proven track record in designing and maintaining end-to-end high-traffic LLM serving systems • Strong problem-solving skills • Ability to communicate technical ideas clearly • Experience with Trusted Execution Environments (TEE) preferred as a nice-to-have • Active contribution to open-source LLM inference engines preferred as a nice-to-have

🏖️ Benefits

• Interview accommodations available upon request

Apply Now

Similar Jobs

🕒 July 28

Olsson

1001 - 5000

🏗️ Construction

Senior Engineer on Railroad Bridge team designing and delivering structural projects. Collaborating with project managers and clients while mentoring junior engineers and ensuring quality solutions.

🕒 July 28

Wand AI

51 - 200

🤖 Artificial Intelligence

🏢 Enterprise

☁️ SaaS

Forward Deployed Engineer focused on designing and deploying AI solutions for enterprises. Involves hands-on technical work and strong customer engagement in a fast-paced startup environment.

🕒 July 28

Redpanda Data

51 - 200

🏢 Enterprise

☁️ SaaS

🤝 B2B

Build production-quality agents and data pipelines on-site with customers as a Forward Deployed Engineer at Redpanda. Collaborate closely with customer teams to integrate solutions effectively.

🇺🇸 United States – Remote

💵 $216.8k - $255k / year

💰 $100M Series D - Redpanda Data on 2025-04

⏰ Full Time

🟡 Mid-level

🟠 Senior

👷🏻‍♀️ Engineer

🕒 July 28

Surge AI

51 - 200

🤖 Artificial Intelligence

🔌 API

☁️ SaaS

Lead Adversarial Engineer managing red teaming against frontier models and ensuring awareness of risks. Collaborating across teams to validate findings and improve security protocols in AI systems.

🕒 July 28

YA Group

501 - 1000

💼 Consulting

🏗️ Construction

🛡️ Insurance

Forensic Engineer investigating damage to various properties and collaborating with clients at YA Group. Requires PE license and 5+ years experience in forensic engineering.

🇺🇸 United States – Remote

💵 $80k - $275k / year

💰 Private equity on 2021-09

⏰ Full Time

🟡 Mid-level

🟠 Senior

👷🏻‍♀️ Engineer