Senior DevOps / MLOps Engineer

Job not on LinkedIn

🔥 0 minutes ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of SimpliGov

SimpliGov

11 - 50 employees

🏛️ Government

☁️ SaaS

⚡ Productivity

Government • SaaS • Productivity

SimpliGov is a powerful end-to-end solution platform that offers workflow process automation, digital forms, and electronic signatures, primarily designed for government applications. The platform features no-code tools, enabling users to automate processes without requiring technical expertise. Its comprehensive tools include document automation, electronic signatures (SimpliSign), a form designer, language access for translations, and performance analytics. Recognized for its applicability in state and local government sectors, SimpliGov facilitates the digitization and modernization of public services, such as DMV services, legal operations, permits, and more. It integrates seamlessly with enterprise systems, ensuring enhanced efficiency and transparency.

📋 Description

• Deploy and operate SimpliGov’s Azure platform, including AKS, networking, identity, storage, and development-through-production environments • Own infrastructure as code end to end, ensuring reproducible environments, drift detection, and platform visibility for all changes • Operate the AI infrastructure layer, including self-hosted observability and evaluation tooling, product telemetry, model gateway, per-workload routing, and compliant GovCloud inference paths • Own cloud and AI cost metering, budgets, unit economics, MACC drawdown strategy, and remediation • Harden production access and controls through least privilege, secrets management, audit evidence, and a FedRAMP-conscious security posture • Partner with AI Operations on deployment and release workflows using Octopus Deploy, environment promotion, progressive rollout, and rollback • Build platform reliability through monitoring, alerting, incident response, and capacity planning • Provide microservices decomposition with service infrastructure, scaling patterns, and clean environment boundaries • Participate in an AI-native product development lifecycle involving autonomous agents in planning, coding, validation, and release • Own and explain everything shipped and surface uncertainty early

🎯 Requirements

• 5+ years in DevOps, platform engineering, or site reliability engineering in SaaS environments • Deep Azure experience, including AKS, networking, Entra identity, and monitoring • Experience running production Kubernetes • Infrastructure as code as the default, using Terraform, Bicep, or similar • Strong scripting skills • MLOps experience deploying and operating LLM or ML systems in production • Experience with model gateways, inference infrastructure, or AI observability stacks • Demonstrated cloud cost analysis and reduction experience • Experience in compliance-heavy environments such as FedRAMP, StateRAMP, SOC 2, or similar is a strong plus • Comfortable holding production access with appropriate discipline • Applicants must not require work visa sponsorship • Must meet US employment authorization requirements through Form I-9/E-Verify

🏖️ Benefits

• Medical, dental, and vision insurance plans, with significant employer contributions for employees AND dependents • Buyup insurance plans available at additional costs • Company-sponsored life/disabilities insurances • 11 Paid holidays • Flexible time off • 401k plan with 4% employer match • Monthly stipends for wellness and home office expenses • Bonus • Benefits

Apply Now

Similar Jobs

🔥 19 minutes ago

VetsEZ

201 - 500

🏥 Healthcare

💼 Consulting

📦 Logistics

Senior Backend DevOps Engineer operating AWS containerized microservices for the VA’s JLV clinical data viewer. Building CI/CD, observability, security, and disaster recovery capabilities.

🔥 57 minutes ago

Group 1001

501 - 1000

💼 Consulting

🏥 Healthcare

💸 Finance

Senior Network Reliability Engineer automating insurance company network platforms at Group 1001. Applying SRE, cloud, Kubernetes, security, and observability practices to improve reliability and reduce operational toil.

🔥 4 hours ago

PathAI

501 - 1000

🏥 Healthcare

💼 Consulting

📦 Logistics

Senior/Staff SRE designing and operating secure on-premises and hybrid-cloud data centers for PathAI’s AI-powered pathology platform. Improving reliability, automation, observability, and incident response for machine-learning infrastructure.

🔥 4 hours ago

Ardent

51 - 200

💼 Consulting

🎖️ Defense

📦 Logistics

DevSecOps Engineer securing cloud products and services for Ardent’s federal national security and defense missions. Automating deployments, vulnerability mitigation, and enterprise system architecture.

🔥 5 hours ago

Hexion Inc.

1001 - 5000

🚘 Automotive

🏗️ Construction

🏭 Manufacturing

Reliability Engineer improving asset performance across Hexion’s North American manufacturing plants. Leading failure elimination, maintenance optimization, and cross-site reliability standardization.