Staff Software Engineer, AI Inference

Job not on LinkedIn

πŸ”₯ 5 minutes ago

Apply Now
Find Similar Remote Jobs

πŸ“Š Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Syllo

Syllo

51 - 200 employees

Founded 2019

βš–οΈ Legal

πŸ€– Artificial Intelligence

☁️ SaaS

πŸ”₯ Funding within the last year

πŸ’° $30M Venture Round - Syllo on 2025-10

Legal β€’ Artificial Intelligence β€’ SaaS

Syllo is a Litigation AI company that provides a unified platform for managing complex litigation end-to-end. Its software combines case management, eDiscovery, legal research, analysis, drafting, and agentic issue tagging to help in-house teams and law firms master the record, apply issue codes at scale, and accelerate strategy and workflow across written discovery, depositions, and pretrial work. Syllo positions itself as an AI-enabled, collaborative workspace focused on high-stakes commercial litigation.

πŸ“‹ Description

β€’ Lead the design and development of our production inference platform. β€’ Define the technical roadmap for inference infrastructure, model serving, and runtime optimization. β€’ Build and operate scalable, cost-effective systems for serving large language models in production. β€’ Evaluate and integrate modern inference technologies, frameworks, and serving runtimes. β€’ Optimize latency, throughput, GPU utilization, memory efficiency, and infrastructure cost. β€’ Develop systems for model deployment, traffic routing, autoscaling, scheduling, observability, and operational excellence. β€’ Partner with ML engineers to productionize new models and inference techniques. β€’ Establish benchmarking methodologies to evaluate new models, runtimes, and hardware. β€’ Make key architectural decisions around when to build internally versus leverage open-source or commercial solutions. β€’ Mentor engineers as the team grows and help establish engineering best practices for AI infrastructure.

🎯 Requirements

β€’ Significant experience designing and operating production AI inference systems. β€’ Experience building or leading production LLM serving infrastructure. β€’ Deep experience with one or more modern inference runtimes and frameworks such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, Hugging Face TGI, NVIDIA Dynamo, or comparable technologies. β€’ Strong background in distributed systems, backend infrastructure, or high-performance platform engineering. β€’ Experience optimizing inference performance across GPU workloads, including latency, throughput, batching, memory utilization, and serving efficiency. β€’ Experience operating GPU infrastructure in production. β€’ Strong proficiency in Python and at least one systems programming language (such as Go, Rust, or C++). β€’ Proven ability to lead technical architecture for complex infrastructure initiatives. β€’ Excellent communication skills and the ability to influence technical direction across engineering teams.

πŸ–οΈ Benefits

β€’ health insurance

Apply Now

Similar Jobs

πŸ”₯ 35 minutes ago

E Source

201 - 500

⚑ Energy

πŸ’Ό Consulting

🀝 B2B

AI / ML Engineer developing scalable ML and AI systems for utilities-focused projects. Collaborating with dynamic teams to implement innovative solutions in a rapidly changing landscape.

AWS

Azure

Cloud

Google Cloud Platform

Python

πŸ”₯ 1 hour ago

Tiger Analytics

1001 - 5000

πŸ₯ Healthcare

πŸ“¦ Logistics

πŸ“£ Marketing

Enterprise AI Architect designing and implementing enterprise-scale AI solutions at Tiger Analytics. Collaborating with business and engineering leaders to drive AI adoption across the organization.

AWS

Azure

Cloud

Kubernetes

Python

πŸ”₯ 2 hours ago

Phoenix Ecommerce Technologies

11 - 50

πŸ›οΈ eCommerce

πŸ’³ Fintech

☁️ SaaS

Staff Engineer developing AI-native commerce solutions for DTC Founders. Building technical foundations for a next-generation storefront platform and collaborating with cross-functional teams.

πŸ‡ΊπŸ‡Έ United States – Remote

πŸ”₯ Funding within the last year

πŸ’° $5M Series unknown on 2025-11

⏰ Full Time

πŸ”΄ Lead

πŸ€– AI Engineer

Angular

JavaScript

Node.js

React

TypeScript

πŸ”₯ 5 hours ago

YipitData

201 - 500

πŸ’Έ Finance

🏒 Enterprise

Lead the development of AI-enabled products and analytics experiences at YipitData. Work on corporate insights and agent systems for major retailers using advanced technologies.

AWS

Python

React

πŸ•’ 2 days ago

Teladoc Health

5001 - 10000

πŸ₯ Healthcare

πŸ‘₯ B2C

☁️ SaaS

AI Engineer developing scalable generative AI and machine learning solutions for healthcare. Collaborating with teams to design and deploy AI/ML pipelines while ensuring production-grade reliability.

Azure

Distributed Systems

Docker

Flask

Kubernetes

Python

Spark

SQL

Terraform