Senior MLOps, ML Platform Engineer

đŸ”„ 0 minutes ago

🌐 Mali, Poland – Remote

infoinfo

⏰ Full Time

🟠 Senior

đŸ—ïž Platform Engineer

đŸ‘» Ghost score 11%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Sigma Software Group

Sigma Software Group

1001 - 5000 employees

Founded 2002

đŸ’Œ Consulting

đŸ„ Healthcare

🚘 Automotive

Consulting ‱ Healthcare ‱ Automotive

Sigma Software Group is a multinational company, established in 2002, that specializes in providing high-quality software development, graphic design, testing, and support services. The company focuses on delivering solutions across various industries such as automotive, telecommunications, aviation, advertising, gaming, banking, real estate, and healthcare. Sigma Software values professional growth, offers remote work opportunities worldwide, and caters to world-renowned clients like AstraZeneca, Scania, and SAS. The company emphasizes a culture of continuous education, mentorship, and flexible work environments, making it a preferred workplace for IT specialists aiming to work on complex solutions utilizing cutting-edge technologies. Sigma Software is committed to innovative solutions and engineering the future while also contributing to social causes such as charitable work in Ukraine.

📋 Description

‱ Build and maintain ML training orchestration pipelines across hourly, daily, and weekly schedules ‱ Implement retries, backfills, and idempotent execution mechanisms ‱ Design and support model registry workflows including versioning, lineage, evaluation gates, and promotion processes ‱ Develop isolated per-advertiser model environments with namespace and configuration separation ‱ Build scalable refresh pipelines and publishing workflows for serving infrastructure ‱ Implement shadow mode and champion/challenger deployment strategies ‱ Develop monitoring and alerting for ML-specific metrics including feature drift, prediction drift, train/serve skew, and calibration decay ‱ Ensure reproducibility of ML workflows using containerized environments, pinned dependencies, and data snapshots ‱ Monitor training and scoring costs across tenants ‱ Collaborate with DevOps and SRE engineers on CI/CD and infrastructure automation ‱ Prepare operational documentation and platform handover materials

🎯 Requirements

‱ 5+ years of experience in MLOps, ML platform engineering, or infrastructure engineering supporting production ML systems ‱ Strong Python skills and experience building platform-level tooling and automation ‱ Hands-on experience with Kubernetes and Docker ‱ Experience building CI/CD pipelines for ML workloads ‱ Hands-on production experience with MLflow, Kubeflow, Airflow, Argo Workflows, Vertex Pipelines, or similar orchestration and ML lifecycle platforms ‱ Experience with ML platforms and model lifecycle tools such as Vertex AI, MLflow, or Kubeflow ‱ Strong understanding of ML observability including drift detection, train/serve skew monitoring, and incident response ‱ Experience designing or supporting multi-tenant ML systems and isolated model environments ‱ Experience working with cloud platforms, preferably GCP ‱ Experience with infrastructure-as-code tools such as Terraform ‱ Experience with Linux environments ‱ Understanding of the ML lifecycle and productionization processes ‱ Upper-Intermediate English level or higher

đŸ–ïž Benefits

‱ Remote work opportunity ‱ Opportunity to work on cutting-edge ML infrastructure projects ‱ Collaboration with experienced engineers ‱ Opportunity to influence architecture decisions ‱ Long-term strategic engagement

Apply Now