Platform Engineer

🔥 0 minutes ago

🌐 Canada, United States – Remote

infoinfo

⏰ Full Time

🟠 Senior

🔴 Lead

🏗️ Platform Engineer

👻 Ghost score 10%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Quantiphi

Quantiphi

1001 - 5000 employees

Founded 2013

💼 Consulting

🏥 Healthcare

📦 Logistics

💰 Series A on 2019-12

Consulting • Healthcare • Logistics

Quantiphi is a leading AI-first digital engineering company that leverages a decade of industry expertise to empower businesses through scalable, secure, and adaptable AI solutions. By integrating cutting-edge technology with real-world applications, Quantiphi transforms organizations across various sectors including healthcare, finance, education, and retail. Their services span AI applications, data analytics, cloud infrastructure modernization, and custom AI implementations. Quantiphi partners with technology giants like AWS, Google Cloud, NVIDIA, and others to drive AI adoption and deliver transformational opportunities for enterprises.

📋 Description

• Design and implement scalable infrastructure for LLM and GenAI workloads across multi-GPU environments • Perform GPU profiling, benchmarking, and performance optimization for distributed training workloads • Manage and schedule compute-intensive jobs using Slurm-based clusters and OpenShift/Kubernetes environments • Enable and optimize the NVIDIA GPU stack • Collaborate with architects, data science, MLOps, application, and AI solution teams • Deploy models in research and production environments • Build and support GenAI pipelines, including fine-tuning, RAG, multi-modal inferencing, and LLMOps • Develop reusable infrastructure templates with Terraform and Helm • Contribute to internal PoCs and workshops and support client-facing delivery engagements • Develop automation software to improve application and cloud-platform functionality, reliability, availability, and manageability • Drive Infrastructure as Code adoption • Design and build self-service, self-healing, synthetic monitoring, and alerting platforms and tools • Automate development and test processes through CI/CD pipelines using Git, Jenkins, SonarQube, Artifactory, and Docker • Build container hosting platforms using Kubernetes • Introduce cloud technologies and tools to drive business value • Lead client technical discussions on architecture design and troubleshooting and proactively provide solutions • Mentor senior resources and team leads

🎯 Requirements

• 10+ years of experience • Strong experience with Slurm and distributed training environments • Hands-on expertise with Red Hat OpenShift and/or Kubernetes • Deep knowledge of NVIDIA GPU ecosystem: CUDA, cuDNN, NCCL, Nsight, Triton/TensorRT • Strong foundation in Linux systems, performance tuning, and multi-GPU optimization • Experience deploying GenAI workloads, including LLM fine-tuning, RAG pipelines, and multi-modal systems • Familiarity with Infrastructure-as-Code tools such as Terraform and Ansible • Experience with cloud GPU environments: GCP, Azure, AWS, OCI, and/or on-prem GPU clusters • Experience with NVIDIA NIMs, DGX systems, or GPU-accelerated containers • Knowledge of LLMOps frameworks and MLOps integration • Familiarity with vector databases and retrieval systems for RAG architectures • Comfortable working in client-facing environments and collaborating with AI solution teams • Healthcare domain experience is nice to have, including FHIR R4, HL7 v2, SMART on FHIR, EHR integration, HIPAA, CDS Hooks, clinical workflows, and clinical decision support systems

🏖️ Benefits

• Up-skill and discover your potential through cutting-edge technology projects • Work in a research-focused organization with 60+ patents filed • Exposure to AI, ML, data, and cloud technologies • Work with Fortune 500 companies • Opportunities to learn, grow, and interact with colleagues globally • Hybrid work culture

Apply Now

Similar Jobs

🕒 5 days ago

Faire

1001 - 5000

🤝 B2B

🛍️ eCommerce

🛒 Retail

Senior Staff ML Platform Engineer architecting Faire’s wholesale technology platform. Defining scalable Databricks, MLOps, and machine learning infrastructure for independent retailers.

🕒 5 days ago

Faire

1001 - 5000

🤝 B2B

🛍️ eCommerce

🛒 Retail

Staff ML Platform Engineer building scalable training, deployment, and governance infrastructure for Faire’s technology wholesale platform. Enabling data scientists to productionize critical machine-learning models.

🕒 August 10

BMO U.S.

5001 - 10000

🛡️ Insurance

💼 Consulting

📦 Logistics

M365 and Power Platform engineer modernizing BMO’s enterprise banking environment. Governing secure collaboration, automation, Dataverse, identity, and AI-enabled platform services.

🇨🇦 Canada – Remote

💵 $75.9k - $141.9k / year

⏰ Full Time

🟡 Mid-level

🟠 Senior

🏗️ Platform Engineer

🕒 August 7

Midnite

201 - 500

🎲 Gambling

👥 B2C

Senior Platform Engineer strengthening backend reliability, service boundaries, and regional scalability for Midnite’s sports betting and gaming platform. Supporting the Canada launch and future market expansion.

🕒 July 31

Orchestry

11 - 50

☁️ SaaS

🔐 Security

🤝 B2B

Senior Platform Engineer developing infrastructure for SaaS product at Orchestry. Responsible for building production-grade, containerized systems and ensuring scalability and reliability in the platform.

🇨🇦 Canada – Remote

💵 $145k - $198k / year

💰 Seed on 2022-12

⏰ Full Time

🟠 Senior

🏗️ Platform Engineer