Search Remote Jobs

ML Cloud Infrastructure Engineer

🔥 13 hours ago

🇺🇸 United States – Remote

💵 $150k - $175k / year

⏰ Full Time

🟡 Mid-level

🟠 Senior

👷 Infrastructure Engineer

👻 Ghost score 20%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of HavocAI

HavocAI

11 - 50 employees

Founded 2024

📦 Logistics

🏭 Manufacturing

🎖️ Defense

💰 Seed Round on 2024-09

Logistics • Manufacturing • Defense

HavocAI is a developer of collaborative autonomy for maritime operations, offering a modular software and vehicle stack that enables fleets of autonomous maritime systems to perform contested logistics, sensor fusion and tracking, domain awareness, and escort-and-engage missions. Their product suite includes onboard autonomy (HAVOC OS), scalable communications (HAVOC CLOUD), and a handheld operator interface (HAVOC CONTROL), marketed as a single solution for theater-scaled security and rapid deployment. HavocAI emphasizes real-time, team-led autonomous solutions that run across diverse environments and supports both hardware (autonomous vessels) and software deployment.

📋 Description

• Build pipelines transforming raw multi-modal data into curated, versioned training datasets • Develop reproducible training and evaluation workflows across cloud compute and GPU resources • Build and maintain model deployment infrastructure for packaging, serving, inference, versioning, and rollback • Implement experiment tracking, dataset lineage, model versioning, and reproducible ML development capabilities • Own data schema versioning and migration across pipelines, data lakes, and services • Design, build, and operate scalable AWS infrastructure using Infrastructure as Code • Build and maintain Kubernetes/EKS workloads and containerized environments • Develop self-service tooling and paved paths for compute scheduling, storage, data access, training, and deployment • Improve utilization, scalability, and cost efficiency across cloud and accelerator infrastructure • Build evaluation frameworks and regression testing for model quality, dataset integrity, and pipeline correctness • Develop monitoring, logging, tracing, and observability across training jobs, data pipelines, and deployed models • Diagnose and resolve performance, scaling, reliability, and infrastructure bottlenecks • Maintain automation, testing, documentation, and operational readiness • Partner with Autonomy, Software, Data, Simulation, and Security teams on ML infrastructure from edge data capture through cloud training and model deployment • Contribute to CI/CD and release processes for models, datasets, and ML pipelines • Translate engineering requirements into scalable platform capabilities • Implement secure infrastructure practices including IAM least privilege, secrets management, access controls, and secure handling of sensitive and defense-related data • Partner with security and infrastructure teams to meet operational and compliance requirements • Build reliable and reproducible pipelines, scalable training and evaluation, self-service ML infrastructure, and improved observability, reliability, security, and cost efficiency within the first 12 months

🎯 Requirements

• 3+ years of experience in software engineering, infrastructure engineering, data engineering, ML infrastructure, or a related field • Strong programming experience in Python; experience in Go, C++, or another systems-oriented language preferred • Experience building and operating production services, APIs, data pipelines, developer platforms, or infrastructure • Hands-on experience with ML workflows such as dataset preparation, model training, evaluation, or deployment • Experience with cloud infrastructure, preferably AWS, and Infrastructure as Code • Hands-on experience with Kubernetes and containerized environments • Strong understanding of production engineering fundamentals, including reliability, observability, testing, automation, and maintainability • Ability to work effectively across engineering disciplines and solve ambiguous technical problems with a high degree of ownership • U.S. Citizenship • Ability to obtain and maintain a U.S. Government security clearance • MLOps and workflow platform experience preferred, including MLflow, Weights & Biases, Kubeflow, Ray, Airflow, or Dagster • GPU/accelerator scheduling, distributed training, or large-scale ML workload experience preferred • Experience with multi-modal datasets including imagery, video, telemetry, sensor, or simulation data preferred • Experience supporting autonomy, robotics, simulation, or real-time systems preferred • Experience deploying ML models to edge or embedded environments preferred • Experience with AWS GovCloud, GCP Assured Workloads, FedRAMP, or IL4/IL5 preferred

🏖️ Benefits

• 100% Employer paid Health, Dental and Vision Insurance for you and your families • Life Insurance (Employer Paid) • Ability to participate in the companies 401k program (Matching) • Unlimited PTO policy with an enforced 2 week minimum • Equity Package • Work / Home Office Stipend • Global Entry • 16 Week Paid Parental Leave • Monthly Health and Wellness Stipend • Bonus offered

Apply Now

Similar Jobs

🔥 18 hours ago

QuickNode ⚡

51 - 200

₿ Crypto

🌐 Web 3

☁️ SaaS

Senior Infrastructure Engineer architecting hybrid cloud, bare-metal, and containerized systems. Quicknode powers global Web3 applications with scalable blockchain infrastructure.

🕒 3 days ago

crewAI

1 - 10

💼 Consulting

📦 Logistics

📣 Marketing

Infrastructure engineer operating AWS, Kubernetes, CI/CD, and observability for CrewAI’s multi-agent AI platform. Improving reliability, security, and self-hosted customer deployments.

🕒 3 days ago

Mirantis

501 - 1000

💼 Consulting

🏥 Healthcare

📦 Logistics

AI storage engineer integrating high-performance NFS with Kubernetes GPU platforms at Mirantis, a Kubernetes-native AI infrastructure company. Automating, tuning, and observing storage across hybrid, edge, and air-gapped deployments.

🕒 3 days ago

Mirantis

501 - 1000

💼 Consulting

🏥 Healthcare

📦 Logistics

Senior engineer integrating and tuning NFS storage for Mirantis’ Kubernetes-native AI infrastructure. Automating observability across GPU workloads, hybrid, edge, and air-gapped environments.

🕒 4 days ago

Delinea

1001 - 5000

🔒 Cybersecurity

☁️ SaaS

🏢 Enterprise

Senior Cloud Engineer securing Delinea’s cloud-native identity security SaaS infrastructure. Designing Azure, Kubernetes, IaC, and compliance solutions for FedRAMP-ready global products.