
11 - 50 employees
Founded 2010
đ¤ Artificial Intelligence
đź Consulting
đ¤ B2B
Artificial Intelligence ⢠Consulting ⢠B2B
Ontrac Solutions is an AI-first technology services firm that helps organizations design, implement, and scale AI-driven systems, cloud architectures, and enterprise platform integrations. They provide AI strategy and generative AI implementation (including RAG and intelligent search), cloud migration and optimization (multi-cloud architecture, Kubernetes, Terraform, GitOps, and FinOps), data and integration engineering, CRM/CMS/e-commerce platform development, and embedded technical talent / staff augmentation to move AI from experimentation into production. With innovation hubs in Chicago and Karachi, Ontrac focuses on delivering production-ready AI and cloud solutions that drive measurable business outcomes.
đ August 9
đ˛đŹ Madagascar â Remote
âł Contract/Temporary
đĄ Mid-level
đ Senior
â DevOps & Site Reliability Engineer (SRE)
đť Ghost score 53%
Ansible
AWS
Cloud
Docker
HAProxy
Java
JavaScript
Kubernetes
Linux
NGINX
Node.js
Prometheus
Puppet
Python
Terraform
Go
Improve your chances of getting an interview by checking your resume score before you apply.

11 - 50 employees
Founded 2010
đ¤ Artificial Intelligence
đź Consulting
đ¤ B2B
Artificial Intelligence ⢠Consulting ⢠B2B
Ontrac Solutions is an AI-first technology services firm that helps organizations design, implement, and scale AI-driven systems, cloud architectures, and enterprise platform integrations. They provide AI strategy and generative AI implementation (including RAG and intelligent search), cloud migration and optimization (multi-cloud architecture, Kubernetes, Terraform, GitOps, and FinOps), data and integration engineering, CRM/CMS/e-commerce platform development, and embedded technical talent / staff augmentation to move AI from experimentation into production. With innovation hubs in Chicago and Karachi, Ontrac focuses on delivering production-ready AI and cloud solutions that drive measurable business outcomes.
⢠Respond to production availability incidents through an on-call rotation ⢠Support service engineers with customer incidents ⢠Use on-call shifts to prevent recurring incidents ⢠Run infrastructure with Ansible, Puppet, Terraform, and Kubernetes ⢠Create symptom-based monitoring and alerting ⢠Document actions and turn findings into repeatable actions and automation ⢠Improve deployment processes ⢠Design, build, and maintain infrastructure scaling to hundreds of thousands of concurrent users ⢠Debug production issues across services and stack levels ⢠Plan infrastructure growth ⢠Code infrastructure automation with Ansible and Terraform ⢠Improve Prometheus monitoring and build new metrics ⢠Help release managers deploy and fix application software versions ⢠Plan and execute migration from AWS virtual machines to Kubernetes-based cloud-native deployments on EKS ⢠Develop relationships with product groups and define SRE KPIs
⢠Cloud-first and security-first mindset ⢠Systems thinking covering edge cases, failure modes, behaviors, and implementations ⢠Familiarity with Linux and Windows ⢠Knowledge of configuration-management systems such as Ansible or Puppet ⢠Strong programming skills in Python, Java, Golang, or Node.js ⢠Ability to collaborate and communicate asynchronously and document work ⢠Proactive approach to fixing broken systems ⢠Experience with Nginx, HAProxy, Docker, Kubernetes, Terraform, or similar technologies
Apply Now