Senior DevOps Engineer – AI Cloud

Job not on LinkedIn

🔥 0 minutes ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Bitdeer Group

Bitdeer Group

201 - 500 employees

💼 Consulting

📦 Logistics

🏗️ Construction

💰 Post-IPO Equity on 2023-05

Consulting • Logistics • Construction

Bitdeer Group is a leader in the blockchain and high-performance computing industry. It is one of the world’s largest holders and suppliers of hash rate, offering specialized mining infrastructure and high-quality hash rate sharing products. Founded by cryptocurrency pioneer Jihan Wu and led by CEO Matt Linghui Kong, the company is headquartered in Singapore with mining datacenters in the United States, Norway, and Bhutan. Bitdeer is committed to providing comprehensive computing solutions, including cloud services and AI capabilities, while emphasizing dedication, authenticity, and trustworthiness in its mission to be the most reliable provider in the industry.

📋 Description

• Design, implement, and maintain end-to-end CI/CD pipelines for software applications and machine learning models • Automate build, testing, deployment, and rollback processes • Build, optimize, and scale cloud-native infrastructure using Kubernetes and Docker • Manage and provision specialized computing resources, including GPU clusters, for AI workloads and model inferencing • Own high-availability design in production environments • Implement disaster recovery, self-healing mechanisms, capacity planning, and performance tuning • Champion Infrastructure as Code practices using Terraform, Ansible, and Helm • Architect and refine monitoring, logging, and alerting systems • Collaborate with R&D, Data Science, Security, and Business teams on workflow optimization and Platform Engineering initiatives • Establish and enforce system stability and security standards, release workflows, Zero Trust access controls, secrets management, and compliance • Lead troubleshooting during complex anomalies and major incidents, conduct root cause analysis, and implement preventative remediation plans

🎯 Requirements

• Bachelor's degree or above in Computer Science, Engineering, or a related technical field • 5+ years of hands-on experience in DevOps, Site Reliability Engineering, or Cloud Infrastructure roles • Expert-level knowledge of Linux operating systems and core networking principles, including TCP/IP, DNS, HTTP, load balancing, and VPCs • Deep mastery of Docker and Kubernetes orchestration, cluster management, and production best practices • Proficiency designing and managing infrastructure on major public or hybrid cloud platforms, including AWS, GCP, Azure, or Alibaba Cloud • Experience with multi-cloud and hybrid-cloud strategies • Strong coding and scripting capabilities in at least one major language, such as Go, Python, or Shell • Practical understanding of CI/CD, Infrastructure as Code, observability, and Site Reliability Engineering principles • Exceptional problem-solving abilities, technical judgment, and cross-team communication skills • a• Preferred: familiarity with MLOps, model serving/inferencing frameworks, GPU clusters, large-scale distributed systems, Internal Developer Platforms, Zero Trust, DevSecOps, SOC2, ISO27001, and technical leadership

Apply Now

Similar Jobs

🔥 27 minutes ago

Shield AI

501 - 1000

🤖 Artificial Intelligence

🚀 Aerospace

🎖️ Defense

Sr. Staff Engineer operationalizing Databricks for Shield AI, a defense-tech company building autonomous aircraft and intelligent systems. Ensuring secure, scalable, observable production data infrastructure.

🔥 4 hours ago

Sonatype

501 - 1000

🔒 Cybersecurity

☁️ SaaS

GCP DevOps Engineer designing secure infrastructure, CI/CD, and Kubernetes platforms for Sonatype. Improving developer delivery, reliability, observability, and software supply-chain security.

🔥 4 hours ago

Summit

51 - 200

💼 Consulting

🏨 Hospitality

📣 Marketing

Site Reliability Architect designing observability and reliability platforms for Summit’s regulated-industry application hosting and cloud services. Improving resilience, automation, and incident response across teams.

🔥 5 hours ago

MeridianLink

501 - 1000

💳 Fintech

🏦 Banking

☁️ SaaS

Senior Site Reliability Engineer operating MeridianLink’s serverless AWS platform. Managing production reliability, databases, backups, monitoring, incident response, and infrastructure automation.

🔥 7 hours ago

NVIDIA

10,000+ employees

🏥 Healthcare

🏭 Manufacturing

🤖 Artificial Intelligence

Senior SRE improving NVIDIA GeForce NOW’s reliable GPU cloud gaming infrastructure. Building observability, automation, Kubernetes, and incident-response tooling for service SLOs.