Machine Learning Scientist

Job not on LinkedIn

🔥 1 hour ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Rime

Rime

11 - 50 employees

🤖 Artificial Intelligence

☁️ SaaS

🤝 B2B

Artificial Intelligence • SaaS • B2B

Rime is a company focused on web-based voice and linguistics technology. The provided site CSS and class names indicate a product that exposes multiple voice personas (professional, casual, formal, energetic), language/multilingual support, and interactive demos and visualizations (globe, languages check). Rime appears to offer a customizable voice/speech platform or interface — likely delivered as an online product aimed at other organizations or developers.

📋 Description

• Design, train, and evaluate speech synthesis models, autoregressive and non-autoregressive. • Drive research on full-duplex and half-duplex multi-modal architectures, including unified S2S systems. • Choose and iterate on speech representations: neural codecs, semantic tokens, mel features, continuous latents. • Build rigorous evaluation, objective and perceptual. Hold the bar on quality and prosodic control. • Collaborate with our linguists on TTS frontend behavior so modeling and frontend choices reinforce each other.

🎯 Requirements

• Deep familiarity with the speech synthesis literature, contemporary and historical — Tacotron, FastSpeech, VITS, VALL-E, the codec-LM lineage. Opinions on what worked and why. • Hands-on training with neural codecs (EnCodec, DAC, Mimi, etc.) and multiple representation choices. • Experience with full- or half-duplex multi-modal modeling (Moshi, LLaMA-Omni, streaming S2S). • Strong attention to detail on data quality. You notice when an annotation pipeline is silently degrading or when an eval set has leakage. • Willing to roll up your sleeves on unglamorous data and training work — paired with the agency to build pipelines so the team isn't stuck doing it by hand. • Working knowledge of TTS frontend (G2P, normalization, prosody) and experience working with linguists. • Strong PyTorch fundamentals. Comfortable with training loops, distributed training, model internals. • PhD or equivalent research experience in speech, audio, ML, or computational linguistics or a track record that makes the credential irrelevant. • Multilingual TTS experience. • Background in prosody or paralinguistics. • Published work in speech, audio, or core ML venues. • Experience taking research models to production: quantization, distillation, streaming inference.

🏖️ Benefits

• Competitive base + meaningful early-stage equity • Remote-friendly • Visa sponsorship available • Access to a proprietary, full-duplex, studio-quality conversational speech corpus • Compute and tooling to do the work • Direct influence on the future of voice AI

Apply Now

Similar Jobs

🔥 6 hours ago

interVal

11 - 50

☁️ SaaS

💳 Fintech

🤖 Artificial Intelligence

Machine Learning Engineer at Interval solving technical problems related to AI, privacy, and distributed systems. Collaborating to build models using enterprise data while maintaining privacy.

🔥 8 hours ago

Adobe

10,000+ employees

💼 Consulting

📣 Marketing

Software Engineer integrating AI/ML models into Adobe's imaging products, working with ML researchers and production teams. Delivering cutting-edge AI features for photographers worldwide.

🔥 9 hours ago

FloatMe

11 - 50

🏥 Healthcare

💼 Consulting

📣 Marketing

Machine Learning Engineer building and evolving ML systems for underwriting at FloatMe. Involved in the full modeling lifecycle to optimize decision-making processes for credit risk.

🔥 11 hours ago

Netflix

10,000+ employees

📱 Media

👥 B2C

Machine Learning Scientist leading research and development of LLMs and VLMs at Netflix. Collaborating with Games Studio R&D team on algorithmic optimization and model adaptation.

🔥 11 hours ago

Humble Robotics

11 - 50

📦 Logistics

🚗 Transport

🚘 Automotive

ML engineer designing and training the vision-language-action foundation model for autonomous driving. Building a production model with a small team tackling a significant challenge in ground transportation.