Site Reliability Engineer

Job not on LinkedIn

🔥 0 minutes ago

🇲🇬 Madagascar – Remote

⏳ Contract/Temporary

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Ontrac Solutions

Ontrac Solutions

11 - 50 employees

Founded 2010

🤖 Artificial Intelligence

💼 Consulting

🤝 B2B

Artificial Intelligence • Consulting • B2B

Ontrac Solutions is an AI-first technology services firm that helps organizations design, implement, and scale AI-driven systems, cloud architectures, and enterprise platform integrations. They provide AI strategy and generative AI implementation (including RAG and intelligent search), cloud migration and optimization (multi-cloud architecture, Kubernetes, Terraform, GitOps, and FinOps), data and integration engineering, CRM/CMS/e-commerce platform development, and embedded technical talent / staff augmentation to move AI from experimentation into production. With innovation hubs in Chicago and Karachi, Ontrac focuses on delivering production-ready AI and cloud solutions that drive measurable business outcomes.

📋 Description

• Participate in an on-call rotation responding to production availability incidents and support service engineers with customer incidents • Use on-call shifts to prevent incidents from recurring • Operate infrastructure with Ansible, Puppet, Terraform, and Kubernetes • Build monitoring and alerting focused on symptoms rather than outages • Document actions and convert findings into repeatable procedures and automation • Improve the deployment process • Design, build, and maintain core infrastructure scaling to hundreds of thousands of concurrent users • Debug production issues across services and stack levels • Plan infrastructure growth • Code infrastructure automation with Ansible and Terraform • Improve Prometheus monitoring and build new metrics • Help release managers deploy and fix new application versions • Plan and execute migration from AWS virtual machines to cloud-native, container-based Kubernetes deployments on EKS • Develop relationships with product groups and define SRE KPIs

🎯 Requirements

• Think cloud-first across public cloud environments • Think security-first • Understand systems, edge cases, failure modes, behaviors, and specific implementations • Familiarity with Linux and Windows • Knowledge of configuration-management systems such as Ansible or Puppet • Strong programming skills in Python, Java, Golang, or Node.js • Ability to collaborate and communicate asynchronously and document work • Go-for-it attitude and willingness to fix broken systems • Experience with Nginx, HAProxy, Docker, Kubernetes, Terraform, or similar technologies

Apply Now