Machine Learning Engineer – Pre Training

Job not on LinkedIn

🕒 May 26

🇺🇸 United States – Remote

💵 $150k - $190k / year

⏰ Full Time

🟢 Junior

🟡 Mid-level

🤖 Machine Learning Engineer

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Mindbeam AI

Mindbeam AI

1 - 10 employees

Founded 2024

💼 Consulting

🏭 Manufacturing

🤖 Artificial Intelligence

Consulting • Manufacturing • Artificial Intelligence

Mindbeam AI is a company building Litespark, an ultra-fast LLM pretraining framework that accelerates training and inference for generative AI applications. Its software improves throughput on existing GPU hardware with zero code changes, is compatible with industry-standard ML frameworks like PyTorch, and claims to reduce pretraining time from months to days while lowering energy consumption and costs (up to ~81% energy savings in cited workloads). Mindbeam targets enterprise and B2B customers seeking scalable, efficient AI infrastructure for large language model development.

📋 Description

• Build scalable pre-training pipelines for foundation models, optimizing throughput and efficiency. • Implement distributed training strategies across GPUs/TPUs and high-performance clusters. • Collaborate with researchers to translate experimental setups into production-ready workflows. • Develop monitoring and fault-tolerance systems to ensure reliable large-scale training. • Continuously benchmark and tune performance across hardware and software stacks.

🎯 Requirements

• Bachelor’s, Master’s, or PhD in Computer Science, Engineering, or related field—or equivalent experience. • 2+ years of experience with large-scale model training and distributed systems. • Strong coding skills in Python and familiarity with ML frameworks (PyTorch, TensorFlow, JAX). • Experience with GPU scheduling, memory optimization, and parallelism strategies. • Comfort with containerized and orchestrated environments (Docker/Kubernetes). • Understanding of high-performance computing and networking bottlenecks.

Apply Now

Similar Jobs

🕒 May 25

Affirm

1001 - 5000

💳 Fintech

👥 B2C

🛍️ eCommerce

Machine Learning Engineer optimizing and automating financial customer operations at Affirm. Collaborating across teams to develop and deploy AI systems for dispute and chargeback handling.

🇺🇸 United States – Remote

💵 $142k - $210k / year

💰 Post-IPO Equity on 2021-01

⏰ Full Time

🟢 Junior

🟡 Mid-level

🤖 Machine Learning Engineer

🦅 H1B Visa Sponsor

info

Airflow

Python

🕒 May 17

Paramount

10,000+ employees

💼 Consulting

📣 Marketing

📱 Media

Machine Learning Engineer optimizing onboarding and re-entry experiences within Paramount's streaming services. Focusing on user commitment and personalization through advanced machine learning techniques.

🇺🇸 United States – Remote

💵 $124k - $186k / year

⏰ Full Time

🟡 Mid-level

🟠 Senior

🤖 Machine Learning Engineer

🦅 H1B Visa Sponsor

info

🕒 May 14

Foundation EGI

11 - 50

🏭 Manufacturing

💼 Consulting

📦 Logistics

ML Ops Engineer at a startup utilizing AI, physics simulation, and computer graphics. Focusing on reducing costs and improving engineering productivity in design and manufacturing.

🇺🇸 United States – Remote

⏰ Full Time

🟡 Mid-level

🟠 Senior

🤖 Machine Learning Engineer

🕒 May 8

Virtusa

10,000+ employees

💼 Consulting

🏢 Enterprise

🤖 Artificial Intelligence

Virtusa AI/ML Engineer designing and deploying Agentic AI and GenAI services remotely. Building Python APIs, agent workflows, Docker solutions, and Azure cloud integrations.

🇺🇸 United States – Remote

💵 zł20.1k - zł23.3k / month

💰 $108M Post-IPO Equity - Virtusa on 2017-05

⏰ Full Time

🟡 Mid-level

🟠 Senior

🤖 Machine Learning Engineer

🕒 May 6

Cambium Learning Group

501 - 1000

📚 Education

🤖 Artificial Intelligence

Machine Learning Engineer II at Cambium Learning Group developing and deploying production-ready ML solutions. Collaborating with cross-functional teams to optimize and transition models into high-performance systems.

AWS

Docker

Flash

Java

Python

PyTorch

Scikit-Learn

Tensorflow