Lead Software Engineer, Model Serving Platform

Likely ghost job

🕒 April 12

🏢🏡 San Francisco – Hybrid

💵 $230k - $300k / year

⏰ Full Time

🟠 Senior

🧑‍💻 Full-stack Engineer

👻 Ghost score 65%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Sciforium

Sciforium

WebsiteLinkedIn

11 - 50 employees

Founded 2024

🤖 Artificial Intelligence

🔌 API

🔧 Hardware

🔥 Funding within the last year

💰 $12M Seed Round - Sciforium on 2025-10

Artificial Intelligence • API • Hardware

Sciforium is a serverless AI infrastructure platform that provides production-ready, multimodal AI services via a unified, OpenAI-compatible API. The company runs vertically integrated AMD GPU hardware and offers model hosting, a model library, real-time evaluation pipelines, and managed agent deployments to help teams build, evaluate, and ship text, image, video, and audio AI applications with lower cost, stronger privacy, and predictable performance.

📋 Description

• Lead the technical direction of the model serving platform, owning architecture decisions and guiding engineering execution. • Build core serving components including execution runtimes, batching, scheduling, and distributed inference systems. • Develop high-performance C++ and CUDA/HIP modules, including custom GPU kernels and memory-optimized runtimes. • Collaborate with ML researchers to productionize new multimodal models and ensure low-latency, scalable inference. • Build Python APIs and services that expose model capabilities to downstream applications. • Mentor and support other engineers through code reviews, design discussions, and hands-on technical guidance. • Drive performance profiling, benchmarking, and observability across the inference stack. • Ensure high reliability and maintainability through testing, monitoring, and engineering best practices. • Troubleshoot and resolve complex issues across GPU, runtime, and service layers.

🎯 Requirements

• Bachelor’s degree in Computer Science, Computer Engineering, Electrical Engineering, or equivalent practical experience • 5+ years of experience designing and building scalable, reliable backend systems or distributed infrastructure • Strong understanding of LLM inference mechanics (prefill vs decode, batching, KV cache) • Experience with Kubernetes/Ray, Containerization • Strong proficiency in C++, Python • Strong debugging, profiling, and performance optimization skills at the system level • Ability to collaborate closely with ML researchers and translate model or runtime requirements into production-grade systems • Effective communication skills and the ability to lead technical discussions, mentor engineers, and drive engineering quality • Comfortable working from the office and contributing to a fast-moving, high-ownership team culture.

🏖️ Benefits

• Medical, dental, and vision insurance • 401k plan • Daily lunch, snacks, and beverages • Flexible time off • Competitive salary and equity

Apply Now

Similar Jobs

🕒 April 12

Koah

1 - 10

📣 Marketing

🤖 Artificial Intelligence

☁️ SaaS

WebsiteLinkedIn

Software Engineer at Koah Labs, shaping AI-native product development and engineering organization with a focus on cross-functional collaboration.

🏢🏡 San Francisco – Hybrid

💵 $180k - $250k / year

🔥 Funding within the last year

💰 $5M Seed on 2025-10

⏰ Full Time

🟡 Mid-level

🟠 Senior

🧑‍💻 Full-stack Engineer

🕒 April 8

Ambrook

11 - 50

🌾 Agriculture

💳 Fintech

🏭 Manufacturing

WebsiteLinkedIn

Software Engineer creating AI applications to improve financial infrastructure for American family-run businesses. Collaborating on experiments and building innovative tools in a fast-paced startup environment.

🏢🏡 San Francisco – Hybrid

💰 Series A on 2022-05

⏰ Full Time

🟡 Mid-level

🟠 Senior

🧑‍💻 Full-stack Engineer

🕒 April 8

Collective

51 - 200

💳 Fintech

☁️ SaaS

WebsiteLinkedIn

Senior Software Engineer at Collective focusing on core product experiences and leveraging AI tools for development. Collaborating across teams to create scalable solutions for members.

🕒 April 8

Persona

51 - 200

💼 Consulting

🏥 Healthcare

📦 Logistics

WebsiteLinkedIn

Software Engineer shaping performance and scalability at Persona. Partnering with product teams to solve critical challenges and mentoring engineers across the organization.

🕒 April 8

Persona

51 - 200

💼 Consulting

🏥 Healthcare

📦 Logistics

WebsiteLinkedIn

Senior Software Engineer developing resilience practices for a configurable identity platform. Collaborating with product teams to enhance performance and scalability in complex systems.