LLM Inference Deployment Engineer

Vaga não está no LinkedIn

🕒 Maio 21

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $180.000 - $240.000 / ano

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🦅 Patrocina Visto H1B

info

🗣️🇺🇸🇬🇧 Inglês obrigatório

Candidatar-se
Encontrar Vagas Remotas Similares

📊 Verifique sua pontuação de currículo para esta vaga

Melhore suas chances de conseguir uma entrevista verificando sua pontuação de currículo antes de se candidatar.

Logo of EnCharge AI

EnCharge AI

11 - 50 funcionários

Fundada em 2022

🤖 Inteligência Artificial

🔧 Hardware

🤝 B2B

💰 $100.000.000 Series B - EnCharge AI em 2025-02

Artificial Intelligence • Hardware • B2B

A EnCharge AI é uma empresa que desenvolve hardware de computação analógica em memória e software complementar para acelerar cargas de trabalho de IA em dispositivos e da borda para a nuvem. Sua tecnologia inclui o acelerador analógico de IA EN100 e outras formas (chiplets, ASICs, placas PCIe) projetadas para fornecer muito maior eficiência energética, densidade computacional e menor custo total de propriedade para inferência em comparação com GPUs convencionais e aceleradores digitais. A EnCharge enfatiza a sustentabilidade, a privacidade de dados por meio do processamento local e a implantação para clientes empresariais e desenvolvedores que buscam computação de IA eficiente e escalável fora da infraestrutura tradicional de nuvem.

Descrição

• Deploy and optimize LLMs (GPT, LLaMA, Mistral, Falcon, etc.) post-training from libraries like HuggingFace • Utilize inference runtimes such as ONNX Runtime, vLLM for efficient execution. • Optimize batching, caching, and tensor parallelism to improve LLM scalability in real-time applications. • Develop and maintain high-performance inference pipelines using Docker, Kubernetes, and other inference servers.

🎯 Requisitos

• Bachelor’s or Master’s degree in Computer Science, Electrical Engineering, or related field. • Experience in LLM inference deployment, model optimization, and runtime engineering. • Strong expertise in LLM inference frameworks (PyTorch, ONNX Runtime, vLLM, TensorRT-LLM, DeepSpeed). • In-depth knowledge of the Python programming language for model integration and performance tuning. • Strong understanding of high-level model representations and experience implementing framework-level optimizations for Generative AI use cases • Experience with containerized AI deployments (Docker, Kubernetes, Triton Inference Server, TensorFlow Serving, TorchServe). • Strong knowledge of LLM memory optimization strategies for long-context applications. • Experience with real-time LLM applications (chatbots, code generation, retrieval-augmented generation).

Candidatar-se

Vagas Similares

🕒 Maio 20

IEX

51 - 200

💸 Finanças

💳 Fintech

🤝 B2B

Systems Reliability Engineer ensuring reliable operations and automation of IEX's trading platform systems. Collaborating with engineering to optimize performance and troubleshoot complex issues.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $150.000 - $225.000 / ano

💰 Corporate Round em 2022-04

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🦅 Patrocina Visto H1B

info

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Maio 20

SouthState Bank

1001 - 5000

🏦 Bancário

💸 Finanças

💳 Fintech

Payment Platform DevOps Engineer at SouthState enabling secure and scalable delivery of cloud-based payment solutions. Collaborating with internal teams for innovation in payment technology.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $152.630 - $243.812 / ano

⏰ Tempo Integral

🟠 Sênior

🔴 Especialista

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Maio 20

LI-COR

201 - 500

🍽️ Alimentos e Bebidas

🏥 Saúde

📦 Logística

Senior DevOps Engineer architecting and managing cloud infrastructure for LI-COR's global IoT platforms. Focus on high-availability operations in the US and China.

🇺🇸 Estados Unidos – Remoto (EUA)

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Maio 18

Tiger Analytics

1001 - 5000

🏥 Saúde

📦 Logística

📣 Marketing

Forward Deployment Engineer integrating and scaling Generative AI solutions collaboratively with clients. Working closely with engineering teams to operationalize AI models across multi-cloud environments.

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Maio 18

decircle

1 - 10

📣 Marketing

📦 Logística

💼 Consultoria

DevOps Engineer for M0, a stablecoin platform optimizing AWS infrastructure and CI/CD pipelines. Collaborating with product teams and ensuring security and performance of cloud-native applications.

🇺🇸 Estados Unidos – Remoto (EUA)

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório