Staff Software Engineer, AI Inference Gateway

Job not on LinkedIn

🔥 0 minutes ago

🇺🇸 United States – Remote

💵 $180k - $220k / year

⏰ Full Time

🔴 Lead

🤖 AI Engineer

🦅 H1B Visa Sponsor

infoinfo

👻 Ghost score 0%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Salad

Salad

11 - 50 employees

Founded 2018

🤖 Artificial Intelligence

💳 Fintech

💰 $17M Series A on 2022-03

Artificial Intelligence • Cloud Computing • Fintech

Salad is a cloud computing company that specializes in democratizing access to GPU resources for AI and machine learning applications. It offers SaladCloud, a platform that connects idle consumer GPUs to businesses in need of computational power, allowing for significant cost savings on cloud services. Salad enables users to deploy AI/ML models at scale efficiently and affordably, leveraging a vast network of distributed nodes worldwide.

📋 Description

• Own the AI Gateway system end to end, including request routing, streaming, batch/async job handling, and horizontal scaling across multiple gateway servers • Build and tune fleet-efficiency algorithms for node autoscaling, performance scoring, and eviction of underperformers • Operate and scale inference nodes running vLLM, llama.cpp, and similar servers • Design and maintain OpenTelemetry-based observability across the gateway and fleet • Maintain accurate pay-per-token and subscription billing integrations based on telemetry • Benchmark new models, quantizations, and GPU hardware for approximately 25% of working time • Work with product and marketing teams to decide what enters production • Collaborate on cross-system design and integrations with Salad engineering • Explain technical trade-offs to technical and non-technical teammates • Report directly to the CTO • Participate in shared after-hours support

🎯 Requirements

• Strong production Rust experience, ideally on high-throughput networked services • Hands-on experience running LLM inference servers (vLLM, llama.cpp, TGI, TensorRT-LLM, or similar) • Knowledge of KV cache behavior • Distributed systems fundamentals: load balancing, backpressure, failure handling, and unreliable nodes • Experience operating systems with metrics, tracing, and on-call responsibilities • Willingness to be on-call • Clear written and verbal communication • Nice to have: Pingora, Tokio, or proxy/gateway internals • Nice to have: experience with heterogeneous or consumer-grade GPU fleets, CUDA, or ROCm • Nice to have: quantization formats including GGUF, AWQ, GPTQ, and FP8 • Strong in at least two of Rust, LLM inference servers, and distributed systems

🏖️ Benefits

• Unlimited PTO • 75% of health insurance premiums covered for you and your dependents • Dental and vision coverage • 401(k) plan • Stock options • Company-provided computer • $500 WFH budget • Fully remote, with flexible hours

Apply Now

Similar Jobs

🔥 9 hours ago

Mom's Meals | A Purfoods Company

501 - 1000

🏥 Healthcare

📦 Logistics

💼 Consulting

Principal AI Engineer advising Mom's Meals leaders and building secure, scalable AI, automation, and integration solutions. Leading responsible AI governance and adoption across the enterprise.

🔥 11 hours ago

dentsu Austria

51 - 200

📣 Marketing

📱 Media

🏢 Enterprise

AI Engineer building production multi-agent systems, LLM orchestration, and data-model interfaces for dentsu’s media business. Delivering scalable client-facing AI automation with Python and cloud platforms.

🕒 Yesterday

CVS Health

10,000+ employees

🏥 Healthcare

⚕️ Healthcare Insurance

🛒 Retail

Staff AI Engineer building secure generative AI, RAG, and agentic systems for CVS Health’s healthcare ecosystem. Leading AWS, GCP, and machine learning architecture in a HIPAA-regulated environment.

🕒 Yesterday

Patterson Companies, Inc.

5001 - 10000

🏥 Healthcare

📦 Logistics

🏭 Manufacturing

Enterprise Data & AI Architect advancing Patterson's enterprise data architecture, governance, integrations, and production AI solutions. Leading automation, data platforms, and cross-functional technical strategy.

🕒 Yesterday

BMO

10,000+ employees

🏦 Banking

💸 Finance

BMO Director architecting AI platforms, cloud systems, and Responsible AI for B2C banking. Driving enterprise adoption across customer experience, personalization, and digital channels.

🇺🇸 United States – Remote

💵 $130k - $270k / year

🔥 Funding within the last year

💰 $142.9M Post IPO debt on 2025-11

⏰ Full Time

🔴 Lead

🤖 AI Engineer