Search Remote Jobs

Senior Machine Learning Engineer

Job not on LinkedIn

🔥 14 hours ago

🔔 Pennsylvania, Virginia – Remote

infoinfo

⏰ Full Time

🟠 Senior

🤖 Machine Learning Engineer

👻 Ghost score 12%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Air

Air

51 - 200 employees

Founded 2017

☁️ SaaS

📱 Media

📣 Marketing

💰 $35M Series B - Air on 2025-01

SaaS • Media • Marketing

Air is an AI-native creative operations and digital asset management (DAM) platform that helps teams organize, find, edit, approve, and multiply creative assets at scale. It combines conversational search, auto-tagging/creative intelligence, AI design and bulk editing (Canvas), reviews & approvals, workflow and brand management, integrations and enterprise security to centralize creative libraries and accelerate content production across channels. Air targets marketing, media, agencies, retail/e‑commerce and product teams seeking faster creative workflows and data-driven asset performance.

📋 Description

• Build infrastructure powering development, evaluation, deployment, and continuous improvement of language models and AI systems • Own LLMOps, fine-tuning infrastructure, model evaluation, dataset pipelines, experiment management, model serving, and production observability • Build platforms enabling AI engineers and researchers to experiment rapidly while maintaining reproducibility, scalability, and reliability • Support the full model lifecycle from dataset creation and experimentation through training, evaluation, deployment, monitoring, and iteration • Design and build LLMOps infrastructure for production language models • Build scalable training and fine-tuning infrastructure for commercial and open-weight language models • Develop pipelines for supervised fine-tuning, parameter-efficient fine-tuning, preference optimization, and other post-training techniques • Build distributed training and GPU-accelerated ML infrastructure • Develop data pipelines for training, fine-tuning, evaluation, and synthetic data generation • Build dataset versioning, lineage, quality validation, transformation, and reproducible experimentation systems • Develop experiment management infrastructure for comparing models, datasets, hyperparameters, prompts, and training techniques • Build automated model evaluation pipelines and production-readiness checks • Design model registries, artifact management, versioning, and promotion workflows • Build and operate scalable model-serving and inference infrastructure • Develop abstractions supporting multiple models and inference providers • Build observability for training and inference, including metrics, tracing, logging, resource utilization, model quality, latency, throughput, and cost • Optimize workloads for GPU utilization, throughput, latency, reliability, and infrastructure cost • Build automated deployment, rollback, canarying, and production-validation workflows • Investigate failures across data pipelines, training jobs, inference services, distributed systems, and production environments • Evaluate emerging models, training techniques, inference frameworks, and ML infrastructure • Partner with AI engineers building agentic systems to provide model, evaluation, and training infrastructure

🎯 Requirements

• U.S. Citizenship is required • 5+ years of experience building production machine learning systems, ML infrastructure, distributed systems, or similar technical systems • Deep experience designing, building, and operating production ML infrastructure or ML platforms • Experience building infrastructure for training, fine-tuning, evaluating, deploying, and monitoring large language models or other large-scale deep learning models • Experience with LLM fine-tuning and post-training workflows, including supervised fine-tuning, LoRA/QLoRA or other parameter-efficient approaches, and preference optimization • Strong understanding of the modern LLM lifecycle, including data preparation, training, evaluation, model artifacts, deployment, inference, monitoring, and iteration • Experience building reproducible ML pipelines involving dataset versioning, experiment tracking, model versioning, and automated evaluation • Experience building and operating production GPU infrastructure across AWS, GCP, Azure, or dedicated GPU providers, including training and/or inference workloads • Strong understanding of distributed systems and computationally intensive ML workloads at scale • Strong programming experience in Python and experience building production-quality software • Deep experience with containers, Kubernetes, and cloud platforms such as AWS, GCP, or Azure • Experience designing scalable APIs, services, asynchronous workloads, and data-processing pipelines • Strong understanding of observability and operational reliability for production ML systems • Comfortable debugging failures across training code, datasets, models, GPUs, distributed systems, and cloud infrastructure • Able to move between ML experimentation and infrastructure engineering • Comfortable working in a rapidly evolving field • Current possession of a U.S. security clearance, or the ability to obtain one with sponsorship (desired) • Experience in or exposure to a startup or entrepreneurial environment (desired) • Experience building secure code execution environments or sandboxes for AI agents (desired) • Experience with multi-agent architectures, agent-to-agent communication, or distributed agent execution (desired) • Experience with fine-tuning, post-training, reinforcement learning, or synthetic data generation (desired) • Experience building AI observability, tracing, and debugging infrastructure (desired) • Experience optimizing inference latency, throughput, GPU utilization, or model-serving costs (desired) • Experience with AI security, adversarial testing, or securing agentic systems (desired) • Experience working in government, defense, or other mission-critical environments (desired)

🏖️ Benefits

• Up to 25% travel, including periodic travel to Pittsburgh, PA and Arlington, VA offices for team collaboration, planning activities, and in-person meetings • Security clearance sponsorship available for candidates able to obtain a U.S. security clearance • Equal Opportunity Employer

Apply Now

Similar Jobs

🔥 16 hours ago

ClickUp

1001 - 5000

☁️ SaaS

⚡ Productivity

🏢 Enterprise

Machine Learning Engineer owning ranking and retrieval ML systems at ClickUp, an AI-powered productivity workspace. Building scalable search relevance, embeddings, and permissions-aware retrieval.

🕒 4 days ago

Spotify

5001 - 10000

📱 Media

👥 B2C

🛍️ eCommerce

Senior ML Engineer building production ML pipelines and LLM evaluation systems for Spotify’s generative AI music products. Creating artist-first listening experiences for fans and musicians.

🕒 5 days ago

Kard

51 - 200

💼 Consulting

📣 Marketing

📦 Logistics

Senior Machine Learning Engineer building production ML systems for Kard’s commerce media infrastructure. Owning personalization, recommendations, model serving, and reusable ML platforms.

🕒 6 days ago

Apella

11 - 50

🏥 Healthcare

🤖 Artificial Intelligence

⚕️ Healthcare Insurance

Senior ML Engineer building MLOps and forecasting infrastructure for Apella’s surgical-care technology. Automating retraining, deployment, monitoring, and scalable production ML systems.

🕒 September 1

Collectors

1001 - 5000

💼 Consulting

📣 Marketing

🛍️ eCommerce

Senior Machine Learning Engineer building production computer vision, multimodal, and agentic AI systems for Collectors’ collectibles platform. Delivering reliable ML tools for graders, researchers, and collectors.