Search Remote Jobs

DevOps, GitLab-based Platform, CICD 30*3 Pipelines

🔥 12 hours ago

🏄 California, Texas – Remote

infoinfo

💵 $180k - $260k / year

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

👻 Ghost score 5%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Bitdeer Group

Bitdeer Group

201 - 500 employees

💼 Consulting

📦 Logistics

🏗️ Construction

💰 Post-IPO Equity on 2023-05

Consulting • Logistics • Construction

Bitdeer Group is a leader in the blockchain and high-performance computing industry. It is one of the world’s largest holders and suppliers of hash rate, offering specialized mining infrastructure and high-quality hash rate sharing products. Founded by cryptocurrency pioneer Jihan Wu and led by CEO Matt Linghui Kong, the company is headquartered in Singapore with mining datacenters in the United States, Norway, and Bhutan. Bitdeer is committed to providing comprehensive computing solutions, including cloud services and AI capabilities, while emphasizing dedication, authenticity, and trustworthiness in its mission to be the most reliable provider in the industry.

📋 Description

• Design, implement, and maintain end-to-end CI/CD pipelines for software applications and machine learning models • Automate build, test, deployment, and rollback processes • Build, optimize, and scale cloud-native infrastructure using Kubernetes and Docker • Manage and provision specialized computing resources, including GPU clusters, for high-performance AI workloads and model inferencing • Own high-availability design in production environments • Implement disaster recovery strategies, self-healing mechanisms, capacity planning, and performance tuning • Champion Infrastructure as Code practices using Terraform, Ansible, and Helm • Architect and refine monitoring, logging, and alerting systems using Prometheus, Grafana, and ELK/EFK stack • Build the Internal Developer Platform and golden paths enabling product, model, and data-science teams to deploy without opening a ticket • Collaborate with R&D, Data Science, Security, and Business teams to streamline workflows and eliminate bottlenecks • Establish and enforce system stability and security standards • Manage release workflows, implement Zero Trust access controls, oversee secrets management, and ensure SOC2/ISO27001 compliance • Lead troubleshooting, root-cause analysis, and preventative remediation during complex anomalies and major incidents • Convert incident learnings into automation to prevent recurrence

🎯 Requirements

• Bachelor's degree or above in Computer Science, Engineering, or a related technical field • 5+ years of hands-on experience in DevOps, Site Reliability Engineering (SRE), or Cloud Infrastructure roles • Expert-level knowledge of Linux operating systems and core networking principles, including TCP/IP, DNS, HTTP, Load Balancing, and VPCs • Deep mastery of Docker and Kubernetes orchestration, including cluster management and production-level best practices • Proficiency designing and managing infrastructure on major public or hybrid cloud platforms, including AWS, GCP, Azure, or Alibaba Cloud • Experience with multi-cloud and hybrid-cloud strategies • Strong coding and scripting capabilities in at least one major language such as Go, Python, or Shell • Systematic and practical understanding of CI/CD methodologies, Infrastructure as Code (IaC), observability paradigms, and SRE principles • Exceptional problem-solving abilities and sharp technical judgment • Excellent cross-team communication skills • Preferred: Familiarity with MLOps practices, model serving/inferencing frameworks such as vLLM, TGI, or Triton Inference Server • Preferred: Experience managing GPU clusters for AI/ML workloads • Preferred: Experience with large-scale distributed systems or high-concurrency environments • Preferred: Hands-on experience designing and building Internal Developer Platforms (IDP) • Preferred: Familiarity with Zero Trust architecture, automated security testing (DevSecOps), SOC2, or ISO27001 • Preferred: Prior experience as a Technical Lead, mentoring junior engineers, or managing DevOps teams • Preferred: Experience wiring an LLM-driven code/config helper into a pipeline or strong opinions on how to • Must comply with applicable work authorization and equal employment requirements in the relevant country, state, and local jurisdictions

Apply Now

Similar Jobs

🔥 12 hours ago

Bitwarden

51 - 200

🔒 Cybersecurity

☁️ SaaS

🏢 Enterprise

Senior SRE operating Bitwarden Gov’s FedRAMP-compliant cloud infrastructure. Managing reliability, monitoring, incident response, Kubernetes, and security across multi-cloud environments.

🔥 12 hours ago

URUS Group

1001 - 5000

🌾 Agriculture

🤝 B2B

🧬 Biotechnology

DevOps Team Lead operating VAS’s AWS platform for farm management software. Leading IaC, reliability, security, cost optimization, and globally distributed workloads.

🔥 14 hours ago

VetsEZ

201 - 500

🏥 Healthcare

💼 Consulting

📦 Logistics

Release Train Engineer leading SAFe delivery and DevOps systems for VetsEZ’s federal healthcare IT project. Overseeing Agile Release Train execution, CI/CD, platform operations, and cross-team delivery.

🔥 19 hours ago

System Automation Corporation

51 - 200

💼 Consulting

🏥 Healthcare

⚖️ Legal

Site Reliability Engineer operating Azure infrastructure for System Automation’s regulatory-agency SaaS platform. Automating reliability, observability, CI/CD, security, and incident response.

🕒 Yesterday

Worth AI

11 - 50

💼 Consulting

🛡️ Insurance

🤖 Artificial Intelligence

Senior DevOps Engineer strengthening Worth AI’s cloud infrastructure, Kubernetes, and reliability systems. Automating AWS infrastructure, CI/CD, observability, disaster recovery, and platform resilience.