Software Engineer – Model Performance

Likely ghost job

🕒 March 27

🏢🏡 San Francisco – Hybrid

💵 $180k - $360k / year

⏰ Full Time

🟡 Mid-level

🟠 Senior

🧑‍💻 Full-stack Engineer

🦅 H1B Visa Sponsor

infoinfo

👻 Ghost score 65%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Baseten

Baseten

WebsiteLinkedIn

11 - 50 employees

Founded 2020

🤖 Artificial Intelligence

☁️ SaaS

🏢 Enterprise

💰 $8M Seed Round on 2022-04

Artificial Intelligence • SaaS • Enterprise

Baseten is a company that provides fast, scalable model inference services, designed for performance, security, and a delightful developer experience. They offer tools to streamline the entire development process, enabling high-throughput inference and fast deployment times. Baseten caters to enterprise companies by delivering robust, secure, and scalable model serving solutions, particularly useful for machine learning and AI model deployment. Their solutions allow organizations to efficiently manage model infrastructure while focusing on creating domain-specific models. Baseten supports open-source model packaging and offers autoscaling features to handle varying demand efficiently.

📋 Description

• Implement, refine, and productionize cutting-edge techniques (quantization, speculative decoding, kv cache reuse, chunked prefill and LoRA) for ML model inference and infrastructure. • Deep dive into underlying codebases of TensorRT, PyTorch, TensorRT-LLM, vllm, sglang, CUDA, and other libraries to debug ML performance issues. • Apply and scale optimization techniques across a wide range of ML models, particularly large language models. • Collaborate with a diverse team to design and implement innovative solutions. • Own projects from idea to production.

🎯 Requirements

• Bachelor's, Master's, or Ph.D. degree in Computer Science, Engineering, Mathematics, or related field. • Experience with one or more general-purpose programming languages, such as Python or C++. • Familiarity with LLM optimization techniques (e.g., quantization, speculative decoding, continuous batching). • Strong familiarity with ML libraries, especially PyTorch, TensorRT, or TensorRT-LLM. • Demonstrated interest and experience in LLM’s. • Deep understanding of GPU architecture. • Proficiency in enhancing the performance of software systems, particularly in the context of large language models (LLMs) (Bonus). • Experience with CUDA or similar technologies (Bonus). • Deep understanding of software engineering principles and a proven track record of developing and deploying AI/ML inference solutions (Bonus). • Experience with Docker and Kubernetes (Bonus).

🏖️ Benefits

• Competitive compensation, including meaningful equity. • 100% coverage of medical, dental, and vision insurance for employee and dependents • Generous PTO policy including company wide Winter Break (our offices are closed from Christmas Eve to New Year's Day!) • Paid parental leave • Company-facilitated 401(k) • Exposure to a variety of ML startups, offering unparalleled learning and networking opportunities.

Apply Now

Similar Jobs

🕒 March 20

Robot.com

51 - 200

📦 Logistics

📣 Marketing

🏭 Manufacturing

WebsiteLinkedIn

Full Stack Developer designing and maintaining web applications for robotic delivery network. Collaborating within cross-functional teams to implement software that is scalable and efficient.

🏢🏡 San Francisco – Hybrid

⏰ Full Time

🟢 Junior

🟡 Mid-level

🧑‍💻 Full-stack Engineer

🕒 March 18

Convex

11 - 50

☁️ SaaS

🏢 Enterprise

WebsiteLinkedIn

Software Engineer designing and maintaining platforms and services at Convex. Contributing to user-facing systems while collaborating closely with customers and teams.

🏢🏡 San Francisco – Hybrid

💵 $170k / year

💰 $100k Venture Round on 2022-07

⏰ Full Time

🟡 Mid-level

🟠 Senior

🧑‍💻 Full-stack Engineer

🦅 H1B Visa Sponsor

infoinfo

🕒 March 18

Convex

11 - 50

☁️ SaaS

🏢 Enterprise

WebsiteLinkedIn

Senior Software Engineer designing and maintaining Convex’s global cloud infrastructure. Collaborating with engineering team to establish reliability practices and prioritize projects.

🏢🏡 San Francisco – Hybrid

💵 $200k / year

💰 $100k Venture Round on 2022-07

⏰ Full Time

🟡 Mid-level

🟠 Senior

🧑‍💻 Full-stack Engineer

🦅 H1B Visa Sponsor

infoinfo

🕒 March 18

Bunkerhill Health

11 - 50

🤖 Artificial Intelligence

🏥 Healthcare

WebsiteLinkedIn

AI Product Engineer role at Bunkerhill Health focusing on product feature ownership and collaboration. Enhance healthcare outcomes by improving backend solutions for patient records and workflows.

🏢🏡 San Francisco – Hybrid

💵 $160k - $260k / year

💰 Seed Round on 2020-12

⏰ Full Time

🟢 Junior

🟡 Mid-level

🧑‍💻 Full-stack Engineer

🕒 March 18

Sentra

11 - 50

🏢 Enterprise

☁️ SaaS

🤖 Artificial Intelligence

WebsiteLinkedIn

Senior Backend Software Engineer at Sentra, developing scalable backend/platform systems and LLM agents for growing startups. Taking ownership and adapting quickly in a fast-paced environment.

🏢🏡 San Francisco – Hybrid

💵 $150k - $300k / year

⏰ Full Time

🟠 Senior

🧑‍💻 Full-stack Engineer