LLM Inference Engineer

🕒 Julho 28

🏄 California – Remoto

infoinfo

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

👷🏻‍♀️ Engenheiro

👻 Score fantasma 24%

infoinfo

🗣️🇺🇸🇬🇧 Inglês obrigatório

Candidatar-se
Encontrar Vagas Remotas Similares

📊 Verifique sua pontuação de currículo para esta vaga

Melhore suas chances de conseguir uma entrevista verificando sua pontuação de currículo antes de se candidatar.

Logo of NEAR AI

NEAR AI

11 - 50 funcionários

Fundada em 2017

🤖 Inteligência Artificial

🔒 Cibersegurança

🏢 Corporativo

Artificial Intelligence • Cybersecurity • Enterprise

A NEAR AI é uma empresa de infraestrutura de IA que prioriza a privacidade, executando inferências confidenciais e agentes de IA autônomos para empresas, governos e aplicações de IA. Sua nuvem NEAR AI e o framework de agentes IronClaw executam modelos e fluxos de trabalho dentro de Ambientes de Execução Confiáveis (TEEs) reforçados por hardware, com isolamento criptográfico, acesso zero do operador e atestação assinada por hardware para verificação independente. A NEAR AI oferece uma plataforma pronta para o uso, permitindo o deploy de modelos privados, automação de operações repetitivas, integração com ferramentas internas (Slack, email, Notion) e oferecimento de cargas de trabalho reguladas ou soberanas através de um marketplace de agentes para finanças, jurídico, operações, pesquisa e segurança.

Descrição

• Architect and maintain production high-traffic LLM serving systems • Optimize throughput, latency, and cost for leading open-source LLMs • Push the boundaries of how large language models are served • Contribute to decentralized and confidential machine learning infrastructure enabling user-owned AI • Help build highly scalable and efficient infrastructure for open-source AI at global scale

🎯 Requisitos

• Strong hands-on experience in LLM inference, with expertise debugging and optimizing major inference engines such as SGLang, vLLM, or TensorRT • Deep knowledge of state-of-the-art GPU architectures • Ability to effectively exploit GPU architectures using PyTorch, Triton, CuTe, CUDA, etc. • Proven track record in designing and maintaining end-to-end high-traffic LLM serving systems • Strong problem-solving skills • Ability to communicate technical ideas clearly • Experience with Trusted Execution Environments (TEE) preferred as a nice-to-have • Active contribution to open-source LLM inference engines preferred as a nice-to-have

🏖️ Benefícios

• Interview accommodations available upon request

Candidatar-se

Vagas Similares

🕒 Julho 28

Olsson

1001 - 5000

🏗️ Construção

Senior Engineer on Railroad Bridge team designing and delivering structural projects. Collaborating with project managers and clients while mentoring junior engineers and ensuring quality solutions.

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Julho 28

Intersect Power

51 - 200

⚡ Energia

PV Performance Engineer ensuring utility-scale solar portfolio efficiency through performance testing and engineering solutions. Collaborating with project teams to optimize generation and reliability of solar projects.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $188.500 - $205.400 / ano

💰 $2.400.000.000 Debt Financing em 2022-09

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

👷🏻‍♀️ Engenheiro

🦅 Patrocina Visto H1B

infoinfo

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Julho 28

LSB Industries, Inc.

501 - 1000

🏭 Manufatura

🌾 Agricultura

⚡ Energia

Controls Project Engineer at LSB Industries focused on managing industrial automation projects. Leading modernization and execution of control systems across multidisciplinary teams in a process plant environment.

🇺🇸 Estados Unidos – Remoto (EUA)

💰 $77.000.000 Grant - LSB Industries em 2024-10

⏰ Tempo Integral

🟠 Sênior

👷🏻‍♀️ Engenheiro

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Julho 28

Inductive Automation

201 - 500

☁️ SaaS

🏭 Manufatura

🤝 B2B

Performance Engineer II at Inductive Automation identifying performance issues and validating system under various configurations. Role offers remote, on-site or hybrid opportunities.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $115.000 - $125.000 / ano

⏰ Tempo Integral

🟢 Júnior

🟡 Pleno

👷🏻‍♀️ Engenheiro

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Julho 28

Wand AI

51 - 200

🤖 Inteligência Artificial

🏢 Corporativo

☁️ SaaS

Forward Deployed Engineer focused on designing and deploying AI solutions for enterprises. Involves hands-on technical work and strong customer engagement in a fast-paced startup environment.

🇺🇸 Estados Unidos – Remoto (EUA)

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

👷🏻‍♀️ Engenheiro

🗣️🇺🇸🇬🇧 Inglês obrigatório