Member of Technical Staff - Training Infrastructure Engineer

🕒 Agosto 9, 2025

🇺🇸 Estados Unidos – Remoto (EUA)

⏰ Tempo Integral

🔴 Especialista

👷 Engenheiro de Infraestrutura

🗣️🇺🇸🇬🇧 Inglês obrigatório

Candidatar-se
Encontrar Vagas Remotas Similares

📊 Verifique sua pontuação de currículo para esta vaga

Melhore suas chances de conseguir uma entrevista verificando sua pontuação de currículo antes de se candidatar.

Logo of Liquid AI

Liquid AI

51 - 200 funcionários

Fundada em 2023

🤖 Inteligência Artificial

🤝 B2B

🏢 Corporativo

Artificial Intelligence • B2B • Enterprise

Liquid AI é uma empresa de tecnologia de ponta que se especializa em soluções de inteligência artificial nativas da borda. Seus inovadores Modelos de Fundação Líquida (LFMs) são projetados para oferecer IA eficiente e personalizável para diversos ambientes—desde computação de borda até infraestruturas em nuvem. Ao maximizar a eficiência computacional e aproveitar arquiteturas avançadas de redes neurais, a Liquid AI fornece às empresas soluções de IA flexíveis e poderosas, adaptadas às suas necessidades específicas.

Descrição

• Design and implement high-performance, scalable training infrastructure that efficiently utilizes our GPU clusters for both specialized and large-scale multimodal models • Build robust data loading systems that eliminate I/O bottlenecks and enable training on diverse multimodal datasets • Develop sophisticated checkpointing mechanisms that balance memory constraints with recovery needs across different model scales • Optimize communication patterns between nodes to minimize the overhead of distributed training for long-running experiments • Collaborate with ML engineers to implement new model architectures and training algorithms at scale • Create monitoring and debugging tools to ensure training stability and resource efficiency across our infrastructure

🎯 Requisitos

• You have extensive experience building distributed training infrastructure for language and multimodal models, with hands-on expertise in frameworks like PyTorch Distributed, DeepSpeed, or Megatron-LM • You're passionate about solving complex systems challenges in large-scale model training—from efficient multimodal data loading to sophisticated sharding strategies to robust checkpointing mechanisms • You have a deep understanding of hardware accelerators and networking topologies, with the ability to optimize communication patterns for different parallelism strategies • You're skilled at identifying and resolving performance bottlenecks in training pipelines, whether they occur in data loading, computation, or communication between nodes • You have experience working with diverse data types (text, images, video, audio) and can build data pipelines that handle heterogeneous inputs efficiently • Desired experience: You've implemented custom sharding techniques (tensor/pipeline/data parallelism). • You have experience optimizing data pipelines for multimodal datasets with sophisticated preprocessing requirements. • You've built fault-tolerant checkpointing systems that can handle complex model states while minimizing training interruptions. • You've contributed to open-source training infrastructure projects or frameworks. • You've designed training infrastructure that works efficiently for either specialized models or large multimodal systems.

🏖️ Benefícios

• U.S. EQUAL EMPLOYMENT OPPORTUNITY INFORMATION • Liquid AI provides equal employment opportunities without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, protected veteran status, or disability. • Completion is voluntary and will not subject you to adverse treatment. • Information obtained will be retained in a confidential file and separate from personnel records.

Candidatar-se