LLM Inference Engineer

🕒 Julho 28

🏄 California – Remoto

infoinfo

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

👷🏻‍♀️ Engenheiro

👻 Score fantasma 43%

infoinfo

🗣️🇺🇸🇬🇧 Inglês obrigatório

Candidatar-se
Encontrar Vagas Remotas Similares

📊 Verifique sua pontuação de currículo para esta vaga

Melhore suas chances de conseguir uma entrevista verificando sua pontuação de currículo antes de se candidatar.

Logo of NEAR AI

NEAR AI

11 - 50 funcionários

Fundada em 2017

🤖 Inteligência Artificial

🔒 Cibersegurança

🏢 Corporativo

Artificial Intelligence • Cybersecurity • Enterprise

A NEAR AI é uma empresa de infraestrutura de IA que prioriza a privacidade, executando inferências confidenciais e agentes de IA autônomos para empresas, governos e aplicações de IA. Sua nuvem NEAR AI e o framework de agentes IronClaw executam modelos e fluxos de trabalho dentro de Ambientes de Execução Confiáveis (TEEs) reforçados por hardware, com isolamento criptográfico, acesso zero do operador e atestação assinada por hardware para verificação independente. A NEAR AI oferece uma plataforma pronta para o uso, permitindo o deploy de modelos privados, automação de operações repetitivas, integração com ferramentas internas (Slack, email, Notion) e oferecimento de cargas de trabalho reguladas ou soberanas através de um marketplace de agentes para finanças, jurídico, operações, pesquisa e segurança.

Descrição

• Architect and maintain production high-traffic LLM serving systems • Optimize throughput, latency, and cost for leading open-source LLMs • Push the boundaries of how large language models are served • Contribute to decentralized and confidential machine learning infrastructure enabling user-owned AI • Help build highly scalable and efficient infrastructure for open-source AI at global scale

🎯 Requisitos

• Strong hands-on experience in LLM inference, with expertise debugging and optimizing major inference engines such as SGLang, vLLM, or TensorRT • Deep knowledge of state-of-the-art GPU architectures • Ability to effectively exploit GPU architectures using PyTorch, Triton, CuTe, CUDA, etc. • Proven track record in designing and maintaining end-to-end high-traffic LLM serving systems • Strong problem-solving skills • Ability to communicate technical ideas clearly • Experience with Trusted Execution Environments (TEE) preferred as a nice-to-have • Active contribution to open-source LLM inference engines preferred as a nice-to-have

🏖️ Benefícios

• Interview accommodations available upon request

Candidatar-se

Vagas Similares

🕒 Julho 28

Olsson

1001 - 5000

🏗️ Construção

Senior Engineer on Railroad Bridge team designing and delivering structural projects. Collaborating with project managers and clients while mentoring junior engineers and ensuring quality solutions.

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Julho 28

Wand AI

51 - 200

🤖 Inteligência Artificial

🏢 Corporativo

☁️ SaaS

Forward Deployed Engineer focused on designing and deploying AI solutions for enterprises. Involves hands-on technical work and strong customer engagement in a fast-paced startup environment.

🇺🇸 Estados Unidos – Remoto (EUA)

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

👷🏻‍♀️ Engenheiro

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Julho 28

Surge AI

51 - 200

🤖 Inteligência Artificial

🔌 API

☁️ SaaS

Lead Adversarial Engineer managing red teaming against frontier models and ensuring awareness of risks. Collaborating across teams to validate findings and improve security protocols in AI systems.

🇺🇸 Estados Unidos – Remoto (EUA)

💰 $25.000.000 Series A em 2020-07

⏰ Tempo Integral

🟠 Sênior

👷🏻‍♀️ Engenheiro

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Julho 28

YA Group

501 - 1000

💼 Consultoria

🏗️ Construção

🛡️ Seguros

Senior Traffic Engineer at YA Group providing expert-level forensic consulting in traffic engineering. Evaluating roadway designs and ensuring compliance with engineering standards.

🇺🇸 Estados Unidos – Remoto (EUA)

💰 Private equity em 2021-09

⏰ Tempo Integral

🟠 Sênior

👷🏻‍♀️ Engenheiro

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Julho 28

TRM Labs

201 - 500

₿ Cripto

📋 Conformidade

🤝 B2B

Agent Engineer developing next-generation AI applications at TRM Labs. Focused on building robust AI infrastructures and agentic systems for investigative tasks.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $200.000 - $275.000 / ano

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

👷🏻‍♀️ Engenheiro

🗣️🇺🇸🇬🇧 Inglês obrigatório