LLM Inference Deployment Engineer

Job not on LinkedIn

🕒 May 21

🌐 United States, Canada – Remote

info

💵 $180k - $240k / year

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🦅 H1B Visa Sponsor

info
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of EnCharge AI

EnCharge AI

11 - 50 employees

Founded 2022

🤖 Artificial Intelligence

🔧 Hardware

🤝 B2B

💰 $100M Series B - EnCharge AI on 2025-02

Artificial Intelligence • Hardware • B2B

EnCharge AI is a company that develops analog in-memory computing hardware and complementary software to accelerate on-device and edge-to-cloud AI workloads. Their technology includes the EN100 analog AI accelerator and other form factors (chiplets, ASICs, PCIe cards) designed to deliver much higher energy efficiency, compute density, and lower total cost of ownership for inference compared with conventional GPUs and digital accelerators. EnCharge emphasizes sustainability, data privacy through local processing, and deployment for enterprise and developer customers seeking efficient, scalable AI computation outside traditional cloud infrastructure.

📋 Description

• Deploy and optimize LLMs (GPT, LLaMA, Mistral, Falcon, etc.) post-training from libraries like HuggingFace • Utilize inference runtimes such as ONNX Runtime, vLLM for efficient execution. • Optimize batching, caching, and tensor parallelism to improve LLM scalability in real-time applications. • Develop and maintain high-performance inference pipelines using Docker, Kubernetes, and other inference servers.

🎯 Requirements

• Bachelor’s or Master’s degree in Computer Science, Electrical Engineering, or related field. • Experience in LLM inference deployment, model optimization, and runtime engineering. • Strong expertise in LLM inference frameworks (PyTorch, ONNX Runtime, vLLM, TensorRT-LLM, DeepSpeed). • In-depth knowledge of the Python programming language for model integration and performance tuning. • Strong understanding of high-level model representations and experience implementing framework-level optimizations for Generative AI use cases • Experience with containerized AI deployments (Docker, Kubernetes, Triton Inference Server, TensorFlow Serving, TorchServe). • Strong knowledge of LLM memory optimization strategies for long-context applications. • Experience with real-time LLM applications (chatbots, code generation, retrieval-augmented generation).

Apply Now

Similar Jobs

🕒 May 20

IEX

51 - 200

💸 Finance

💳 Fintech

🤝 B2B

Systems Reliability Engineer ensuring reliable operations and automation of IEX's trading platform systems. Collaborating with engineering to optimize performance and troubleshoot complex issues.

🇺🇸 United States – Remote

💵 $150k - $225k / year

💰 Corporate Round on 2022-04

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🦅 H1B Visa Sponsor

info

🕒 May 20

SouthState Bank

1001 - 5000

🏦 Banking

💸 Finance

💳 Fintech

Payment Platform DevOps Engineer at SouthState enabling secure and scalable delivery of cloud-based payment solutions. Collaborating with internal teams for innovation in payment technology.

🇺🇸 United States – Remote

💵 $152.6k - $243.8k / year

⏰ Full Time

🟠 Senior

🔴 Lead

⛑ DevOps & Site Reliability Engineer (SRE)

🕒 May 18

decircle

1 - 10

📣 Marketing

📦 Logistics

💼 Consulting

DevOps Engineer for M0, a stablecoin platform optimizing AWS infrastructure and CI/CD pipelines. Collaborating with product teams and ensuring security and performance of cloud-native applications.

🇺🇸 United States – Remote

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🕒 May 14

NVIDIA

10,000+ employees

🏥 Healthcare

🏭 Manufacturing

🤖 Artificial Intelligence

Senior Network Reliability Engineer maintaining NVIDIA's cloud and datacenter networks. Engaging in global support and driving operational improvements across teams.

🕒 May 14

Avaya

5001 - 10000

💼 Consulting

📣 Marketing

📦 Logistics

Site Reliability Engineer at Avaya driving stability and performance across Azure and GCP platforms. Collaborating with DevOps and Security teams to manage incidents and optimize operations.

🇺🇸 United States – Remote

💵 $129k - $143k / year

💰 Post-IPO Debt on 2022-06

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🦅 H1B Visa Sponsor

info