Staff Machine Learning Engineer – Platform, MLOPS

🔥 12 hours ago

🚗 Michigan – Remote

infoinfo

💵 $154.1k - $226k / year

⏰ Full Time

🔴 Lead

🏗️ Platform Engineer

🦅 H1B Visa Sponsor

infoinfo

👻 Ghost score 0%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Credit Acceptance

Credit Acceptance

1001 - 5000 employees

Founded 1972

💸 Finance

💳 Fintech

🚘 Automotive

Finance • Fintech • Automotive

Credit Acceptance is a company that specializes in providing auto financing solutions to individuals, including those with bad credit history or no credit history at all. With over 50 years of experience, Credit Acceptance has helped more than 4 million consumers secure financing for vehicle purchases. The company operates a network of over 15,000 dealers and offers customers the ability to pre-qualify for auto financing, which does not impact their credit score. Furthermore, Credit Acceptance provides educational resources and tools, such as payment and auto price calculators, to assist customers in managing their finances and understanding their credit options. The company's mission is to drive possibility by offering financial solutions that give individuals the opportunity to improve their financial situation and secure reliable transportation.

📋 Description

• Own and operate the platform that every ML and AI model at Credit Acceptance runs on • Own the end-to-end deployment path for ML and GenAI models, including training and inference pipelines, model registry and versioning, serving endpoints, and controlled promotion across development, QA, and production • Own runtime health for production models through monitoring, alerting, drift and quality-regression detection, latency and throughput objectives, capacity, autoscaling, and incident response through root cause and corrective action • Operate the agent runtime layer through the enterprise AI Gateway and MCP Gateway; migrate agents onto the governed path and maintain scoped, versioned, least-privilege tool surfaces • Own production evaluation of agents, including online scoring, behavioral monitoring, drift detection, sampling and judge pipelines, and release gates • Build and maintain observability and evaluation infrastructure, including multi-turn and multi-step agent trace and telemetry capture, logging standards, pipeline plumbing, and data contracts • Measure and manage platform unit economics, including cost per inference, document, and interaction, and provide cost analysis for build-versus-buy and hosting decisions • Deliver reusable pipeline templates, deployment patterns, reference implementations, and internal tooling to enable adoption of the standard ML platform path • Partner with Cloud Engineering, Data Engineering, Security, and SRE on enterprise governance, identity, and observability • Respond to AI-specific production incidents such as prompt injection, rogue-agent cost spikes, data-classification exposure, delegation abuse, and model endpoint failures; drive corrective actions to closure • Maintain architecture documentation and system diagrams for the ML platform • Mentor engineers and interns on production ML practices and raise operating standards through design and code review • Work from home with occasional planned travel to the assigned Southfield, Michigan office; may work at that office if requested by the team member • Remain compliant with company policies, processes, and legal guidelines • Perform other duties as assigned; attendance as required by department

🎯 Requirements

• Bachelor's degree in Computer Science, Engineering, Statistics or a relevant technical field with at least 7 years of relevant experience, or a Master's degree in one of those fields with at least 5 years of relevant experience • 5+ years building and operating production ML or AI systems, with direct ownership of at least two of: training or inference pipelines, model serving infrastructure, model registry and versioning, or production monitoring and alerting • Demonstrated ownership of a production ML or AI service through its full operating life • Strong Python and SQL, with production-quality engineering practice including version control, testing, code review, and CI/CD applied to ML workloads • Hands-on experience with a cloud ML platform in production; AWS and Databricks strongly preferred, including model serving, job orchestration, and a model registry or experiment tracking system such as MLflow • Working knowledge of operational differences between LLM/GenAI workloads and traditional ML, including token cost and latency behavior, caching and batching, and non-deterministic output • Experience running LLM or agent applications behind a gateway or proxy layer, including model routing and fallback, credential and key management, rate limiting, and budget enforcement • Working knowledge of tool-calling architecture for agents, including the Model Context Protocol, MCP servers, tool scoping and authorization, and gateway-brokered tool access • Experience with containerization and infrastructure as code • Ability to communicate technical and non-technical trade-offs clearly in writing • Preferred: production experience with GPU-backed model serving, OpenTelemetry and enterprise observability platforms such as Dynatrace, model serving efficiency techniques, Databricks Unity Catalog, regulated-industry AI systems, agentic or multi-step AI systems, managed agent/tool gateways such as AWS Bedrock AgentCore Gateway, governed MCP servers, and emerging agent interoperability and identity standards such as Agent2Agent (A2A) • Ability to communicate complex technical information verbally and in writing to all levels, including senior leadership • Ability to solve problems at the source with simple, working solutions • Prompt and effective incident, task, and project resolution • Demonstrated ability and motivation to teach others • Ability to build relationships across and vertically throughout the organization • Ability to prioritize and execute tasks in a high-pressure environment • Required degrees must have been earned at accredited institutions of higher education recognized by the Council for Higher Education Accreditation or equivalent

🏖️ Benefits

• Annual variable bonus of cash and equity, between 10-20% • 401(K) match • Adoption assistance • Parental leave • Tuition reimbursement • Comprehensive medical, dental, and vision coverage • Many nonstandard benefits • Potential premium on top of the posted range for candidates residing in San Francisco, Seattle, Boston, New York City, Los Angeles, and San Diego zones

Apply Now

Similar Jobs

🔥 13 hours ago

Northrop Grumman

10,000+ employees

🏭 Manufacturing

📦 Logistics

🎖️ Defense

Teamcenter deployment and Kubernetes engineer operating secure PLM environments for Northrop Grumman’s defense technology programs. Automating containerized delivery through Red Hat, OpenShift, and CI/CD pipelines.

🕒 Yesterday

CentralReach

201 - 500

🏥 Healthcare

📚 Education

Principal Platform Engineer defining AWS cloud platforms, reliability, and developer tooling. Supporting CentralReach’s autism and IDD care software used by clinicians, educators, and care teams.

🇺🇸 United States – Remote

💵 $200k - $215k / year

💰 Private equity on 2018-03

⏰ Full Time

🔴 Lead

🏗️ Platform Engineer

🕒 3 days ago

Jones Lang LaSalle Americas, Inc.

10,000+ employees

🏠 Real Estate

🤝 B2B

💼 Consulting

Staff Product Manager owning JLL’s Falcon AI developer platform for real estate engineering. Building agent, MCP, CI/CD, observability, and governance capabilities at enterprise scale.

🕒 6 days ago

EasyLlama - HR & Compliance Training For Modern Teams

11 - 50

💼 Consulting

🏥 Healthcare

🏭 Manufacturing

Founding Platform Engineer building infrastructure for EasyLlama’s AI-powered compliance and human risk management platform. Owning Ruby on Rails, Kubernetes, cloud, billing, integrations, and AI platform engineering.

🕒 6 days ago

Coinbase

1001 - 5000

💼 Consulting

₿ Crypto

💸 Finance

Staff Software Engineer leading React and TypeScript architecture for Coinbase’s financial reconciliation platform. Building auditable workflows that help Finance and Accounting monitor and investigate digital asset transactions.