MLOps, LLMOps Engineer – Mid-Level

🔥 0 minutes ago

🇮🇳 India – Remote

⏰ Full Time

🟡 Mid-level

🟠 Senior

🤖 Machine Learning Engineer

👻 Ghost score 16%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Irth Solutions

Irth Solutions

51 - 200 employees

Founded 1995

☁️ SaaS

⚡ Energy

📡 Telecommunications

💰 Private Equity Round on 2015-07

SaaS • Energy • Telecommunications

Irth Solutions is a market-leading provider of a SaaS platform focused on enhancing resilience and reducing risk in the management of critical network infrastructure. Their solutions are trusted by energy, utility, and telecom companies across the U. S. and Canada, offering capabilities in damage prevention, training, asset inspections, land management, and 811 ticket management. By leveraging business intelligence, analytics, and geospatial data, Irth Solutions provides comprehensive situational awareness to proactively manage and mitigate risks in network infrastructure. The acquisition of OneBridge Solutions enhances their asset performance management offerings.

📋 Description

• Operationalize the complete ML lifecycle—training, evaluation, packaging, deployment, and monitoring—on Databricks. • Implement ML workflows using Bronze → Silver → Gold medallion architecture with Delta Lake. • Establish Unity Catalog model-management patterns for governance, lineage, discovery, and access control. • Develop reusable templates for ML/LLM jobs, workflows, and deployment processes. • Create and maintain cluster policies aligned with enterprise platform guardrails. • Productionize ML and GenAI models for damage prevention, asset integrity, land management, and stakeholder engagement use cases. • Design, build, and maintain production-grade LLM and RAG pipelines. • Implement vector search and retrieval architectures using technologies such as Databricks Vector Search. • Deploy and manage model-serving and inference endpoints; optimize performance, scalability, reliability, and cost. • Implement batch, streaming, and online inference patterns and establish service-level expectations and SLAs. • Integrate data contracts, quality gates, PII controls, data residency policy-as-code, and end-to-end lineage into the ML lifecycle. • Use Databricks Asset Bundles and GitHub Actions to version, test, and promote ML/LLM assets across DEV, QA, and PROD. • Build automated unit, integration, regression, data-quality, and model-quality test suites with deployment gates. • Instrument pipelines and services for SLOs, monitoring, alerting, Jira ticket creation, model/data/feature drift detection, and LLM observability. • Develop runbooks, troubleshooting procedures, operational documentation, disaster-recovery procedures, and participate in on-call rotations and DR testing. • Enforce cost and ownership tags, support showback/chargeback reporting, monitor AI infrastructure and inference costs, and identify optimization opportunities. • Collaborate with Data Science, Data Engineering, Platform, Product, Security, and domain teams.

🎯 Requirements

• 3–5 years of experience in MLOps, LLMOps, ML Engineering, Data Engineering, or platform-focused ML engineering. • Hands-on experience with Databricks Jobs and Workflows, Delta Lake, Unity Catalog, and Databricks SQL Warehouses. • Experience building and maintaining CI/CD pipelines for data and ML workloads using GitHub Actions and Databricks Asset Bundles (DABs). • Experience with DEV → QA → PROD environment promotion and parameterized deployments. • Experience with secure secrets management using Azure Key Vault, AWS KMS/Secrets Manager, or equivalent technologies. • Strong understanding of data contracts, schema governance, and automated data/feature validation. • Experience with Great Expectations-style validation frameworks or equivalent rule-based data-quality solutions. • Experience building observable production pipelines with metrics, dashboards, alerting, and SLO monitoring. • Practical experience with RBAC/ABAC, Unity Catalog security, PII detection and obfuscation, private networking, data-access controls, and policy-as-code for data residency. • Strong proficiency in Python and SQL. • Working knowledge of distributed computing and job orchestration within Databricks/Spark environments. • Ability to troubleshoot production ML/data workloads and participate in operational support and incident resolution. • Preferred: hands-on experience with LLM/GenAI workflows, prompt engineering, RAG, LLM evaluation, AI safety and guardrails, retrieval/response-quality evaluation, latency optimization, and token/API-cost optimization. • Preferred: experience with geospatial data and analytics, including PostGIS, spatial joins, spatial indexing and tiling, coordinate systems and projections, and GIS-based feature engineering. • Preferred: experience integrating Power BI with Databricks SQL Warehouses and semantic layers. • Preferred: practical knowledge of FinOps, including resource tagging, budget management, cost monitoring, showback/chargeback, and cost anomaly detection. • Preferred: knowledge of Databricks disaster-recovery patterns, including Delta Lake Deep Clone, Delta Sharing, cross-region recovery, tiered RTO/RPO strategies, and DR testing. • Preferred: hands-on experience with Microsoft Azure and AWS. • Preferred: understanding of cloud-native security patterns, including Private Link, VPC/VNet connectivity and peering, egress restrictions, KMS, AWS Secrets Manager, Azure Key Vault, and data-plane isolation.

🏖️ Benefits

• Competitive Salary – A competitive compensation package based on experience and qualifications. • Medical, Dental, and Vision Insurance – Comprehensive insurance coverage to support you and your family. • 401(k) Plan with Company Match. • Generous Paid Time Off (PTO) – Time off to support work-life balance and personal needs. • Company-Paid Holidays – Paid holidays throughout the year. • Flexible Work Options – Work-from-home opportunities are available, depending on role and business needs. • On-Call Compensation – Additional pay for eligible on-call shifts.

Apply Now

Similar Jobs

🕒 2 days ago

NewRocket

501 - 1000

💼 Consulting

🏥 Healthcare

🛡️ Insurance

Forward Deployed AI/ML Engineer developing and deploying models, LLMs, and Agentic AI systems for NewRocket’s enterprise ServiceNow clients. Building scalable cloud, data, and automation solutions.

Apache

AWS

Azure

Cloud

ETL

MongoDB

MySQL

Numpy

Pandas

Postgres

Python

PyTorch

Scikit-Learn

Spark

SQL

Tableau

Tensorflow

🕒 August 27

Pythian

201 - 500

💼 Consulting

🏥 Healthcare

📦 Logistics

AI/ML Engineer deploying scalable ML pipelines, LLMs, and Generative AI solutions. Supporting Pythian’s cloud data transformation services across client and internal products.

AWS

Azure

Cloud

Docker

ETL

Google Cloud Platform

Kubernetes

Python

PyTorch

Scikit-Learn

Tensorflow

🕒 July 31

Jumio Corporation

201 - 500

💼 Consulting

🏥 Healthcare

📦 Logistics

Machine Learning Engineer enhancing fraud detection solutions for ID verification at Jumio. Engaging with advanced machine learning, deep learning, and computer vision technologies.

AWS

Cloud

Google Cloud Platform

Python

PyTorch

Scikit-Learn

Tensorflow

🕒 July 31

Jumio Corporation

201 - 500

💼 Consulting

🏥 Healthcare

📦 Logistics

Machine Learning Engineer with expertise in biometrics focusing on face recognition systems. Leading design, development, and optimization of ML models using AWS technologies.

Airflow

AWS

Cloud

EC2

Python

PyTorch

Tensorflow

🕒 July 30

Coinbase

1001 - 5000

💼 Consulting

₿ Crypto

💸 Finance

Machine Learning Engineer at Coinbase working on building a unified orchestration layer for AI experiences. Collaborating with cross-functional teams to enhance customer-facing and internal systems.

🇮🇳 India – Remote

💵 ₹4.4M / year

💰 $21.4M Post-IPO Equity on 2022-11

⏰ Full Time

🟢 Junior

🟡 Mid-level

🤖 Machine Learning Engineer

AWS

Python