ML Ops Engineer

🕒 Abril 20

🇺🇦 Ucrânia – Remoto

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

🤖 Engenheiro de Machine Learning

👻 Score fantasma 56%

infoinfo

🗣️🇺🇸🇬🇧 Inglês obrigatório

Candidatar-se
Encontrar Vagas Remotas Similares

📊 Verifique sua pontuação de currículo para esta vaga

Melhore suas chances de conseguir uma entrevista verificando sua pontuação de currículo antes de se candidatar.

Logo of Pragmatike

Pragmatike

11 - 50 funcionários

Fundada em 2022

💼 Consultoria

📣 Marketing

🎯 Recrutamento

Consulting • Marketing • Recruitment

Pragmatike é uma empresa de recrutamento e seleção de TI remota que busca, avalia e coloca talentos internacionais de tecnologia para empresas. Eles oferecem um processo de correspondência humanizado (não apenas por AI), avaliações técnicas e colocações rápidas, afirmando que especialistas qualificados podem ser apresentados em até 48 horas, enquanto cuidam da integração, faturamento e folha de pagamento internacional. A Pragmatike atende startups e grandes empresas com cargos como desenvolvedores, engenheiros de dados, desenvolvedores mobile e de games, gerentes de produto e especialistas, operando em mais de 60 países com um grande pool de talentos avaliados.

Descrição

• Build and operate production-grade model serving infrastructure using frameworks such as vLLM, TGI, Triton, or equivalent • Design and implement robust deployment pipelines with blue/green and canary rollout strategies for ML models • Develop and maintain auto-scaling systems, multi-model serving architectures, and intelligent request routing layers • Optimize GPU utilization, memory efficiency, network throughput, and model artifact storage performance • Design observability systems for tracking inference latency, throughput, GPU usage, cost metrics, and system health • Manage model registries and CI/CD pipelines enabling automated and reproducible model deployments • Own the full lifecycle of ML systems from development through production, including operational support and on-call responsibilities • Define engineering best practices and contribute to platform scalability in a fast-moving startup environment

🎯 Requisitos

• 4+ years of experience in ML Ops, Platform Engineering, SRE, or similar infrastructure roles focused on ML systems • Hands-on experience with model serving frameworks such as vLLM, TGI, Triton, or equivalent • Strong background in container orchestration and operating GPU-based workloads in production • Experience with MLOps tooling including model registries, experiment tracking, and automated deployment pipelines • Proficiency in Python and infrastructure-as-code tools (e.g., Terraform, Helm, or similar) • Strong understanding of distributed systems, performance tuning, and production reliability engineering • Ability to effectively use AI coding assistants to accelerate development and debugging workflows • Ownership mindset with the ability to operate independently in a remote-first environment.

🏖️ Benefícios

• Take ownership of critical infrastructure powering a rapidly scaling AI-native cloud platform • Build foundational ML inference systems from the ground up in a high-growth, well-funded startup • Work at the intersection of distributed systems, GPU computing, and sustainable cloud architecture • Gain deep expertise in next-generation AI infrastructure and large-scale model serving systems • Influence core engineering decisions and define best practices that will scale with the company.

Candidatar-se

Vagas Similares

🕒 Março 13

Globaldev Group

201 - 500

💼 Consultoria

📦 Logística

☁️ SaaS

Senior Machine Learning Engineer taking end-to-end responsibility for ML initiatives in programmatic advertising. Shaping ML architecture and strategy with potential to lead a team.

🇺🇦 Ucrânia – Remoto

⏰ Tempo Integral

🟠 Sênior

🤖 Engenheiro de Machine Learning

🗣️🇺🇸🇬🇧 Inglês obrigatório