Search Remote Jobs

AWS DevOps Engineer, MLOps

Job not on LinkedIn

🔥 16 hours ago

🗽 New York – Remote

infoinfo

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

👻 Ghost score 10%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Capstone Integrated Solutions

Capstone Integrated Solutions

51 - 200 employees

💼 Consulting

🛒 Retail

Consulting • Retail

Capstone Integrated Solutions is a full-service software and services company that specializes in retail software solutions. The company provides a wide range of IT services, including custom development, application development, project management, and systems integration. With a focus on customer experience and support, Capstone operates on a build-as-a-service model, delivering comprehensive technology solutions for various platforms and industries, particularly in the retail sector.

📋 Description

• Design and maintain CI/CD pipelines supporting automated testing, deployment, and change management across CDP, data lake, marketing cloud, and migration workstreams • Provision and manage AWS environments including S3 buckets, Glue jobs, networking, IAM, and VPCs for non-production development and staging • Ensure environments are logically segregated with no public internet access per SOW requirements • Implement and enforce infrastructure as code using AWS CDK or Terraform • Manage the MLOps lifecycle on Amazon SageMaker, including model versioning, deployment automation, endpoint monitoring, and automated retraining triggers • Support Amazon Bedrock integration deployments by managing prompt versioning, model configuration, and API endpoint reliability • Implement and maintain data lake security infrastructure using AWS Lake Formation, including row-level security, column-level encryption, and IAM permission boundaries • Support Azure-to-AWS migration execution, including target environment provisioning, AWS DataSync jobs, cutover procedures, freeze periods, and rollback plans • Monitor pipeline health, data job execution, and ML endpoint performance across AWS Glue, Step Functions, Lambda, Kinesis, and SageMaker using CloudWatch and AWS-native observability tools • Implement AWS cost optimization practices and provide guidance on Azure resource decommissioning post-migration • Coordinate change management processes, including documenting changes, obtaining approvals, executing peer-reviewed deployments, and maintaining rollback procedures • Collaborate with Data/ML and Full Stack engineers across data pipelines, ML inference, and application layers • Contribute to operations runbooks, deployment guides, and infrastructure documentation

🎯 Requirements

• 4+ years of DevOps or cloud infrastructure experience, with at least 2+ years on AWS • Hands-on experience with CI/CD tooling such as GitHub Actions, AWS CodePipeline, Jenkins, or equivalent in a multi-environment AWS setup • Proficiency with infrastructure as code using AWS CDK, CloudFormation, or Terraform • Working knowledge of MLOps practices on Amazon SageMaker, including model deployment, endpoint management, monitoring, and pipeline automation • Experience with AWS IAM, Lake Formation, VPC, and security best practices for data environments • Familiarity with AWS monitoring and observability tools including CloudWatch, CloudTrail, and AWS Config • Experience supporting data pipeline infrastructure including AWS Glue, Step Functions, Lambda, Kinesis, and S3 • Strong understanding of environment management, change control processes, and rollback procedures in enterprise delivery contexts • Ability to work in Agile/Scrum teams alongside AWS Professional Services with structured change approval workflows • Azure infrastructure and Azure-to-AWS migration experience (nice to have) • Familiarity with Amazon Bedrock deployment and API management (nice to have) • Experience with Amazon Connect or omnichannel platform infrastructure (nice to have) • Knowledge of data compliance and privacy controls (nice to have) • AWS Certification such as DevOps Engineer Professional, Solutions Architect, or Machine Learning Specialty (nice to have) • Exposure to Kiro CLI or AI-assisted infrastructure tooling (nice to have) • Background in real estate, property management, or enterprise SaaS environments (nice to have)

🏖️ Benefits

• Full-time contract employment • Equal opportunity employer committed to diversity and an inclusive, safe environment • Learning and professional growth culture • Knowledge sharing and continuous learning culture

Apply Now

Similar Jobs

🔥 17 hours ago

NVIDIA

10,000+ employees

🏥 Healthcare

🏭 Manufacturing

🤖 Artificial Intelligence

Senior SRE maintaining NVIDIA’s managed DGX Cloud AI clusters across major cloud providers. Improving Kubernetes reliability, observability, GPU workloads, and production incident response.

🔥 18 hours ago

Pluribus Digital

51 - 200

💼 Consulting

🏥 Healthcare

📦 Logistics

Senior DevOps Engineer designing secure, scalable Azure architectures for government agencies. Leading cloud migration, IaC, CI/CD, governance, and federal compliance strategies.

🔥 19 hours ago

FindErnest

11 - 50

💼 Consulting

🏢 Enterprise

🤝 B2B

DevOps Lead managing production Azure AKS environments and advanced Terraform IaC for IT services. Driving Flux CD, Helm, Dynatrace observability, CI/CD, security, and team performance.

🔥 19 hours ago

Guidehouse

10,000+ employees

🏥 Healthcare

🎖️ Defense

📦 Logistics

Cloud infrastructure engineer automating secure AWS and hybrid environments for Guidehouse’s federal clients. Building Terraform, CI/CD, security, monitoring, and disaster recovery capabilities.

🔥 22 hours ago

Qodo (formerly Codium)

11 - 50

🤖 Artificial Intelligence

☁️ SaaS

Senior DevOps Engineer owning Kubernetes-based deployments of Qodo’s AI code review platform in customer AWS, GCP, and Azure environments. Troubleshooting infrastructure and automating enterprise onboarding.