Search Remote Jobs

Lead DevOps/AIOps Engineer

đŸ”„ 0 minutes ago

🩀 Maryland – Remote

info

đŸ’” $130k - $155k / year

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🩅 H1B Visa Sponsor

info
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Blend360

Blend360

501 - 1000 employees

đŸ„ Healthcare

🏹 Hospitality

✈ Travel

💰 $100M Private Equity Round on 2022-08

Healthcare ‱ Hospitality ‱ Travel

Blend360 is a professional services company specializing in AI, data analytics, and data-driven solutions. They work with Fortune 1000 and large enterprise brands to tackle significant challenges by integrating people and artificial intelligence. Blend360 focuses on several domains including business intelligence, data engineering, data science, MLOps, and data governance. Their industries of expertise encompass financial services, energy, healthcare and life sciences, retail, technology, media & telecom, and travel & hospitality. Blend360 is recognized for their AI and data solutions, having earned accolades such as "AI-Enabling Solution of the Year" and being listed among the "Top Generative AI Service Providers 2024.

📋 Description

‱ Lead the design and implementation of cloud-native DevOps and MLOps architectures on GCP ‱ Build and optimize CI/CD pipelines for data, ML, and application workloads ‱ Develop infrastructure-as-code using tools such as Terraform and establish repeatable deployment patterns ‱ Architect and operationalize data platforms using BigQuery, Cloud Storage, Dataflow, Pub/Sub, Dataproc, and Cloud Composer ‱ Build MLOps capabilities across model development, deployment, monitoring, versioning, and retraining ‱ Establish observability across data and ML platforms, including logging, monitoring, alerting, pipeline health, data quality, and model performance ‱ Implement secure, scalable cloud infrastructure using GCP IAM, networking, secrets management, and security controls ‱ Partner with Data Engineers, ML Engineers, Architects, and client stakeholders to translate business requirements into production-ready solutions ‱ Establish engineering standards for deployment automation, testing, environment management, reliability, and operational excellence ‱ Troubleshoot complex production issues and drive root-cause analysis and long-term remediation ‱ Mentor engineers and serve as a technical leader across DevOps, cloud, data, and MLOps initiatives ‱ Evaluate emerging GCP and AI technologies for business and engineering value

🎯 Requirements

‱ 7+ years of experience in DevOps, cloud engineering, platform engineering, MLOps, or a related discipline ‱ Strong hands-on experience with Google Cloud Platform, particularly BigQuery and cloud-native data services ‱ Experience designing and implementing end-to-end data platforms on GCP ‱ Strong understanding of BigQuery architecture, performance optimization, data ingestion, partitioning, clustering, and data security ‱ Experience with CI/CD, Git, automated testing, containerization, and Kubernetes/GKE ‱ Strong Infrastructure-as-Code experience, preferably Terraform ‱ Experience with Vertex AI and/or production ML platforms, including model deployment and monitoring ‱ Experience with Cloud Composer/Airflow, Dataflow, Dataproc/Spark, and Pub/Sub ‱ Strong understanding of observability, reliability engineering, monitoring, logging, and alerting ‱ Proficiency with Python and/or Bash ‱ Strong understanding of cloud security, IAM, networking, secrets management, and enterprise governance ‱ Ability to operate at architectural and hands-on engineering levels ‱ Excellent communication skills and ability to work with technical teams and senior client stakeholders ‱ Nice to have: Vertex AI, MLflow, Kubeflow, GenAI/LLM production solutions, Docker and Kubernetes/GKE, data quality, data lineage, metadata management, semantic data layers, AWS or Azure, and consulting/professional services experience

Apply Now

Similar Jobs

đŸ”„ 17 minutes ago

NVIDIA

10,000+ employees

đŸ„ Healthcare

🏭 Manufacturing

đŸ€– Artificial Intelligence

Senior Site Reliability Engineer operating NVIDIA’s Base Command Manager and large-scale GPU clusters. Handling incidents, deployments, and resilient Slurm/Kubernetes infrastructure for AI data centers.

đŸ”„ 51 minutes ago

SAIC

10,000+ employees

☁ SaaS

📣 Marketing

🏱 Enterprise

DevSecOps Engineer securing SAIC’s Windows, Linux, and cloud infrastructure. Managing Active Directory, automation, CI/CD pipelines, and compliance in controlled defense environments.

đŸ‡ș🇾 United States – Remote

đŸ”„ Funding within the last year

💰 $500M Post-IPO Debt - SAIC on 2025-09

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

đŸ”„ 1 hour ago

Coupa Software

1001 - 5000

đŸ’Œ Consulting

📩 Logistics

đŸ„ Healthcare

Lead Active Directory Site Reliability Engineer architecting secure identity infrastructure for Coupa’s AI-powered spend management platform. Driving automation, cloud integration, observability, and least-privilege administration across global environments.

đŸ”„ 3 hours ago

InnoData

2 - 10

đŸ€ B2B

đŸ’Œ Consulting

🌍 Social Impact

Application Reliability Engineer supporting GCP and Google App Engine applications for Innodata, a global data engineering and AI services company. Managing incidents, deployments, microservices, and reliability improvements.

đŸ”„ 7 hours ago

General Dynamics Information Technology

10,000+ employees

đŸ’Œ Consulting

đŸ„ Healthcare

📩 Logistics

DevSecOps Security Engineer automating AWS cloud security, compliance, and vulnerability workflows for GDIT’s U.S. government missions. Managing continuous ATO pipelines, IaC security, and audit evidence at scale.