AI Infrastructure Engineer III

🕒 August 4

🇪🇬 Egypt – Remote

⏰ Full Time

🟡 Mid-level

🟠 Senior

👷 Infrastructure Engineer

👻 Ghost score 15%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Mozn

Mozn

201 - 500 employees

Founded 2017

💼 Consulting

🏥 Healthcare

📦 Logistics

💰 $10M Series A on 2022-02

Consulting • Healthcare • Logistics

Mozn is a regional AI company building Arabic-native generative AI and enterprise AI platforms. It provides OSOS (an Arabic-first GenAI platform), FOCAL (a financial-crime and fraud detection platform), and customized AI solutions spanning language intelligence, risk intelligence, operational AI, data management, geospatial intelligence, and AI centers. Mozn focuses on serving enterprise customers in the MENA region (including healthcare, finance, and government) with SaaS products and tailored AI services that prioritize cultural relevance, data security, and regulatory compliance.

📋 Description

• Design, deploy, and operate enterprise AI/ML platforms • Build self-service platforms for Data Scientists and ML Engineers • Deploy and operate Kubeflow, MLflow, KServe, Ray, or similar AI platforms • Design infrastructure supporting model training, experimentation, feature engineering, and inference • Build highly available and scalable model serving infrastructure • Design and operate GPU clusters for large-scale AI workloads • Optimize GPU scheduling, utilization, sharing, autoscaling, and resource allocation • Deploy and manage NVIDIA GPU Operator and GPU-enabled Kubernetes environments • Optimize distributed GPU training performance across multi-node clusters • Troubleshoot AI infrastructure performance bottlenecks • Build CI/CD pipelines for ML workloads • Automate AI infrastructure provisioning using Infrastructure as Code • Implement monitoring and observability for GPU utilization, model serving, training jobs, and inference latency • Collaborate with Data Science teams to improve platform usability, performance, and reliability

🎯 Requirements

• 4–6 years of experience in AI Infrastructure, MLOps, Platform Engineering, or Cloud Engineering • Strong hands-on experience with Kubernetes • Experience with Kubeflow, MLflow, or similar ML platform technologies • Experience operating GPU infrastructure for AI workloads • Strong understanding of NVIDIA GPU technologies, CUDA fundamentals, and GPU optimization • Experience supporting distributed training workloads • Experience with model serving platforms such as KServe, Triton Inference Server, Ray Serve, or similar • Experience with AWS, GCP, OCI, or Azure AI platforms • Experience automating infrastructure using Terraform, Helm, GitOps, or Ansible • Strong scripting or programming skills in Python, Bash, or Go • Experience with Prometheus, Grafana, OpenTelemetry, ELK/OpenSearch, or equivalent observability platforms • Preferred: experience with PyTorch, TensorFlow, Hugging Face, or JAX • Preferred: experience with distributed training frameworks such as Ray, DeepSpeed, Horovod, or NCCL • Preferred: experience with vector databases, LLM infrastructure, RAG architectures, or GenAI platforms • Preferred: experience operating inference platforms for large language models • Preferred: experience supporting AI research or Data Science teams in production environments • Preferred: contributions to Cloud Native, Kubernetes, AI, or ML open-source communities • Cloud, Kubernetes, NVIDIA, or AI/ML certifications are a plus

🏖️ Benefits

• Competitive compensation • Top-tier health insurance • Enabling culture • Responsibility and trust • Freedom and autonomy in the role • Fun and dynamic workplace • Opportunity to work alongside leading AI professionals • Inclusive and empowering workplace culture

Apply Now

Similar Jobs

🕒 July 27

Talent 360 ME

51 - 200

💼 Consulting

🎯 Recruiter

🤝 B2B

Senior Engineer responsible for managing and optimizing IT infrastructure and systems at an organization. Ensuring high availability, security, and performance of IT operations remotely.

🇪🇬 Egypt – Remote

💰 Series unknown on 2025-02

⏰ Full Time

🟠 Senior

👷 Infrastructure Engineer

Cloud

Cyber Security

DNS

Linux

Shell Scripting

VMware

🕒 June 24

Müller's Solutions

11 - 50

💼 Consulting

Integration/Infrastructure Specialist supporting enterprise integration and infrastructure initiatives within the ServiceNow ecosystem. Requires expertise in ServiceNow, MID Server, GCP, and SSO technologies.

🇪🇬 Egypt – Remote

⏰ Full Time

🟡 Mid-level

🟠 Senior

👷 Infrastructure Engineer

Cloud

DNS

Firewalls

Google Cloud Platform

JavaScript

Python

ServiceNow

Shell Scripting

SOAP

🕒 October 13, 2025

Unifonic

501 - 1000

💼 Consulting

🏥 Healthcare

📦 Logistics

Senior Infrastructure Engineer in the DevX team at Unifonic, enhancing CI/CD and ensuring high product availability. Work within a collaborative environment on innovative communication solutions.

🇪🇬 Egypt – Remote

💰 $125M Series B on 2021-09

⏰ Full Time

🟠 Senior

👷 Infrastructure Engineer

Ansible

Cloud

Docker

Kubernetes

Oracle

Python

Terraform