Machine Learning Engineer – ML Training Platform

Job not on LinkedIn

🔥 0 minutes ago

🌐 United States, Australia – Remote

infoinfo

⏰ Full Time

🟡 Mid-level

🟠 Senior

🤖 Machine Learning Engineer

👻 Ghost score 25%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Pluralis Research

Pluralis Research

1 - 10 employees

🤖 Artificial Intelligence

🌐 Web 3

Artificial Intelligence • Web 3

Pluralis Research is a foundational AI research lab focused on Protocol Learning — decentralized, multi‑participant training of foundation models where no single participant holds a full copy of the model. The group develops methods to enable communication‑efficient model and pipeline parallelism, unextractable collaborative models, and high‑compression context parallelism so community‑trained, community‑owned frontier models can scale over low‑bandwidth, internet‑connected devices. Their work targets practical systems and algorithms that make decentralized training competitive with centralized training while enabling new ownership and economic models.

📋 Description

• Architect, build, and scale the platform supporting continuous experimentation and large-scale training on consumer nodes and cloud instances • Design resource management systems provisioning and orchestrating compute across AWS, GCP, and Azure using Pulumi/Terraform • Handle dynamic scaling, state synchronization, and concurrent operations across hundreds of heterogeneous nodes • Architect fault-tolerant distributed ML infrastructure for GPU clusters and NVIDIA runtime • Implement S3 checkpointing, large-dataset management and streaming, health monitoring, and resilient retry strategies • Build systems simulating and handling bandwidth shaping, latency injection, and packet loss • Manage node churn and maintain data flow across workers with heterogeneous connectivity

🎯 Requirements

• Production experience with infrastructure-as-code (Pulumi/Terraform/CloudFormation) managing multi-cloud deployments • Experience with Docker/Kubernetes (EKS), GPU workloads, and heterogeneous clusters at scale • Understanding of distributed training workflows, including checkpointing, data sharding, model versioning, and long-running job orchestration • Experience with decentralized networking, including P2P, NAT traversal, traffic shaping, and real bandwidth constraints • Strong Python engineering, including asyncio, concurrency, retry logic, cloud SDKs, and CLI tooling • Hands-on observability and SRE practice, including Prometheus/Grafana, performance profiling, and incident response • Experience in a startup with heavy service orchestration or at big-tech scale • Ability to demonstrate which systems you owned • Belief in Protocol Learning as a viable path for collective, trustless, and sovereign AI • Professional-level English proficiency, written and spoken • Comfortable working across time zones

🏖️ Benefits

• Significant equity/ownership for key technical contributors in addition to a high base salary • Flexible work environment with team members distributed globally • Optional full visa sponsorship and relocation support to either Australia or the US • Opportunity to work on novel, unpublished problems in training and serving frontier models

Apply Now

Similar Jobs

🔥 5 hours ago

CrowdStrike

5001 - 10000

🔒 Cybersecurity

☁️ SaaS

🤖 Artificial Intelligence

Threat Analyst improving CrowdStrike’s machine-learning cybersecurity detections through malware, binary, and false-positive analysis. Supporting internal teams and Data Science in protecting customers from breaches.

🔥 6 hours ago

Odyssey

1001 - 5000

🎖️ Defense

🏥 Healthcare

📦 Logistics

Machine Learning Engineer building secure ML platforms, computer vision systems, and AI microservices. Supporting Odyssey’s defense, ISR, and warfighter readiness capabilities.

🔥 10 hours ago

LouisianaNOW.Jobs

51 - 200

🎯 Recruiter

🏪 Marketplace

AI/ML Architect modernizing enterprise case-management software for GDIT’s U.S. government customers. Building cloud ML pipelines, MLOps infrastructure, and secure AI services.

🔥 11 hours ago

General Dynamics Information Technology

10,000+ employees

💼 Consulting

🏥 Healthcare

📦 Logistics

AI/ML Architect SME modernizing enterprise case-management software for U.S. government missions at GDIT. Designing cloud ML pipelines, MLOps infrastructure, and secure AI services.

🔥 12 hours ago

General Dynamics Information Technology

10,000+ employees

💼 Consulting

🏥 Healthcare

📦 Logistics

AI/ML Architect SME designing cloud MLOps architectures and pipelines for GDIT’s government case-management modernization program. Supporting secure, scalable enterprise applications.