Search Remote Jobs

DevOps Engineer, AI/ML Infrastructure

🔥 18 hours ago

🗽 New York – Remote

infoinfo

💵 $106k - $127k / year

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🦅 H1B Visa Sponsor

infoinfo

👻 Ghost score 0%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of CreatorIQ

CreatorIQ

201 - 500 employees

Founded 2014

📣 Marketing

📱 Media

☁️ SaaS

Marketing • Media • SaaS

CreatorIQ is a leading influencer marketing software platform that assists brands across various industries in growing, managing, scaling, and measuring their creator programs. It provides tools for discovering creators, executing campaigns, and optimizing performance with advanced analytics. CreatorIQ is designed to enhance brand engagement by leveraging the power of content and connections with trusted creators and communities. The platform is used by global enterprises, media and sports organizations, and direct-to-consumer brands to track, manage, and maximize their influencer marketing efforts.

📋 Description

• Support and maintain scalable, highly available, and secure cloud infrastructure • Provision and manage cloud resources using Infrastructure as Code • Implement cloud security best practices, IAM and role-based access controls, encryption, vulnerability management, and secure infrastructure configurations • Support containerized environments and orchestration platforms • Apply DevSecOps principles across infrastructure and deployment workflows • Participate in disaster recovery planning, testing, and recovery activities • Maintain and optimize CI/CD pipelines supporting application and ML model deployments • Improve deployment reliability and support zero-downtime deployment strategies • Automate configuration management, infrastructure provisioning, and routine operational processes • Troubleshoot deployment and pipeline issues and prevent recurrence • Develop scripts and automation to reduce manual work • Design, deploy, operate, and secure infrastructure supporting AI and agentic products • Operate and scale ML platform infrastructure, including Databricks clusters, jobs compute, ML pipelines, and Model Serving endpoints • Manage production model-serving infrastructure, compute capacity, provisioned throughput, and autoscaling • Monitor model drift, data quality, inference performance, and serving health • Maintain monitoring, logging, metrics, and alerting using Prometheus, Grafana, Coralogix, and CloudWatch • Support incident response and perform Root Cause Analysis for infrastructure and deployment issues • Partner with Software Engineers, ML Engineers, QA, Software Engineers in Test, IT Security, and Product Support • Respond to engineering and Product Support requests • Maintain accurate internal technical and operational documentation • Collaborate with international teams across multiple time zones

🎯 Requirements

• 3+ years of experience in DevOps, Cloud Engineering, Site Reliability Engineering (SRE), or a similar infrastructure-focused role • 2+ years of hands-on experience with AWS services such as EC2, S3, RDS, Lambda, IAM, VPC, SQS, and API Gateway, or similar services • 2+ years of experience with containerized environments and orchestration platforms such as Kubernetes and Amazon EKS • Strong experience building and maintaining CI/CD pipelines using GitLab CI/CD or Jenkins • Hands-on experience with Infrastructure as Code using Terraform, Terragrunt, CloudFormation, or similar technologies • Strong Linux system administration and troubleshooting skills • Solid understanding of networking fundamentals, including routing, load balancing, and network security • Scripting experience with Python, Bash, or similar languages • Hands-on experience using AI tools to improve engineering workflows, automation, troubleshooting, or agentic use cases • Experience supporting data, ML, or other compute-intensive production workloads • Experience with Google Cloud would be valuable • Familiarity with Helm and service mesh technologies such as Istio, Linkerd, or Traefik would be beneficial • Experience with serverless and event-driven architectures is a plus • Exposure to cloud and infrastructure security practices, vulnerability management, and tools such as Nessus, Prowler, Trivy, or firewalls would be valuable • Knowledge of security standards, compliance requirements, and cloud security best practices is beneficial • Experience with observability, log analysis, and monitoring platforms such as Coralogix, Prometheus, or Grafana is a plus • FinOps experience would be valuable • Experience with API gateways or API management platforms such as Kong or Apigee is beneficial • Experience with MLOps platforms and practices, particularly Databricks, model serving, ML pipelines, and model monitoring, would be an advantage

🏖️ Benefits

• 15 days of vacation • Floating and company holidays • Wellness benefits • Paid parental leave • Comprehensive medical insurance • Dental insurance • Vision insurance • Life insurance • Disability insurance • 401(k) plan • Work from home stipend • Learning platform, training, and tools • Flexible work model combining in-person and remote work

Apply Now

Similar Jobs

🔥 19 hours ago

Latitude IT Solutions | SDVOSB

11 - 50

🤝 B2B

💼 Consulting

🏛️ Government

Enterprise Architect leading CI/CD, infrastructure, and security architecture for Latitude’s federal health supply-chain software. Driving automated delivery and mentoring engineering teams.

🔥 19 hours ago

GE Vernova

10,000+ employees

💼 Consulting

📦 Logistics

🏭 Manufacturing

Leading SRE architecture, release governance, incident response, and FinOps for GE Vernova’s GridOS utility SaaS platform. Managing a distributed team supporting critical North American utility infrastructure.

🔥 21 hours ago

Millennium

201 - 500

💼 Consulting

🎖️ Defense

🔒 Cybersecurity

Remote DevSecOps cybersecurity engineer securing DoD software, GitLab CI/CD pipelines, containers, and vulnerability management. Supporting Millennium’s national-security cybersecurity missions through RMF compliance and software assurance.

🔥 21 hours ago

GitLab

1001 - 5000

💼 Consulting

📣 Marketing

🤖 Artificial Intelligence

Distinguished Engineer directing GitLab’s CI, CD, Plan, and source-control architecture. Driving scalable AI-native DevOps systems, technical strategy, and engineering quality.

🔥 22 hours ago

Vultr

201 - 500

🤖 Artificial Intelligence

🤝 B2B

🔧 Hardware

Senior SRE maintaining MySQL and PostgreSQL reliability for Vultr’s global cloud infrastructure. Owning monitoring, disaster recovery, incident response, security compliance, and automation.