Machine Learning Platform Engineer

🔥 0 minutes ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of BJAK

BJAK

51 - 200 employees

🛍️ eCommerce

🛡️ Insurance

🏪 Marketplace

eCommerce • Insurance • Marketplace

BJAK is a leading online platform in Southeast Asia that offers comprehensive automobile insurance comparison services. The company enables Malaysian users to compare and purchase auto insurance from multiple insurers efficiently, providing considerable savings and convenience. BJAK is renowned for its user-friendly digital platform that allows quick insurance and road tax renewals, offering discounts up to 11%. With a strong emphasis on customer service, BJAK also provides 24/7 roadside assistance, accident support, and replacement vehicles. It is a pioneer in the insurance comparison sector in the region and has facilitated significant savings for millions of car owners.

📋 Description

• Build and operate the ML infrastructure and platforms powering A1’s AI products • Design systems for model training, evaluation, deployment, inference, and experimentation • Build and optimise model serving and inference infrastructure for high-throughput and low-latency workloads • Improve reliability, scalability, latency, and cost efficiency of AI systems • Develop pipelines for data preparation, training, evaluation, model release, and continuous improvement • Build platforms and tooling that enable AI engineers and researchers to experiment, evaluate, and ship models faster • Develop evaluation and benchmarking infrastructure to measure model quality, performance, and regressions • Build production observability, monitoring, tracing, and alerting for AI/ML workloads • Identify bottlenecks across the ML stack and continuously improve system performance • Collaborate with AI engineers, researchers, and product teams to turn model requirements into production-ready infrastructure • Ensure AI infrastructure reliably supports production workloads at scale • Enable efficient model training, evaluation, deployment, and improvement • Deliver inference systems with strong latency, throughput, reliability, and cost efficiency • Ensure ML pipelines are reproducible, observable, maintainable, and robust • Detect and diagnose model and infrastructure regressions quickly • Create reusable ML infrastructure platform primitives • Enable the AI stack to evolve as new models, architectures, and inference techniques emerge

🎯 Requirements

• Strong software engineering fundamentals and experience building production systems • Experience building ML infrastructure, platforms, or production machine learning systems • Experience with model deployment, inference, evaluation, or data pipelines • Strong understanding of distributed systems and system reliability • Ability to write clean, maintainable, production-quality code • Experience with Python • Experience with PyTorch and/or JAX • Familiarity with LLM/ML serving infrastructure such as vLLM, SGLang, or TensorRT-LLM • Experience with cloud infrastructure • Experience with distributed systems • Experience with ML/data pipelines and workflow orchestration • Experience with GPU infrastructure and performance tooling • Experience with vector databases and retrieval infrastructure • Ability to work in ambiguous, fast-moving environments • Ownership, experimentation, and continuous improvement mindset

Apply Now

Similar Jobs

🕒 6 days ago

Dremio

201 - 500

📚 Education

🤝 B2B

☁️ SaaS

Senior Software Engineer building observability, infrastructure-as-code, and platform services for Dremio’s unified lakehouse platform. Scaling telemetry, distributed systems, cloud operations, and SAP HANA/BDC integrations for enterprise customers.

Apache

AWS

Azure

Cloud

Distributed Systems

DNS

Docker

Google Cloud Platform

Java

JavaScript

Jenkins

Kubernetes

Microservices

MongoDB

Node.js

NoSQL

Python

Redis

Rust

SQL

TCP/IP

Terraform

Go

🕒 July 31

HumanIT Digital Consulting

51 - 200

💼 Consulting

📦 Logistics

📣 Marketing

Platform Engineer responsible for the architecture and governance of Microsoft Fabric and Power Platform. Focus on security, compliance, and operational stability of AI and data platforms.

Azure

Cloud

🕒 July 30

Expleo Group

10,000+ employees

💼 Consulting

🎖️ Defense

📦 Logistics

Platform Engineer responsible for technical architecture of the Microsoft AI and Data Platform with focus on Microsoft Fabric. Expleo delivers engineering, technology, and consulting services globally.

Azure

🕒 June 13

Devoteam

5001 - 10000

💼 Consulting

🏥 Healthcare

📣 Marketing

Senior AWS Platform Engineer designing and building enterprise-scale cloud infrastructure solutions at Devoteam. Join our multidisciplinary team of cloud experts driving efficiency and security across high-availability production environments.

AWS

Cloud

DNS

Docker

EC2

Flux

Grafana

Jenkins

Kubernetes

Prometheus

Python

Terraform

Go

🕒 June 9

Intermedia Cloud Communications

1001 - 5000

💼 Consulting

🏥 Healthcare

⚖️ Legal

Senior Resource Plane Platform Engineer building and evolving shared infrastructure services for cloud communications company. Collaborating in a fast-paced environment with a high-impact team in Portugal.

Cloud

Distributed Systems

ElasticSearch

Kafka

Kubernetes

Linux

RabbitMQ

Redis

Terraform