Site Reliability Engineer

🔥 0 minutes ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Ontrac Solutions

Ontrac Solutions

11 - 50 employees

Founded 2010

🤖 Artificial Intelligence

💼 Consulting

🤝 B2B

Artificial Intelligence • Consulting • B2B

Ontrac Solutions is an AI-first technology services firm that helps organizations design, implement, and scale AI-driven systems, cloud architectures, and enterprise platform integrations. They provide AI strategy and generative AI implementation (including RAG and intelligent search), cloud migration and optimization (multi-cloud architecture, Kubernetes, Terraform, GitOps, and FinOps), data and integration engineering, CRM/CMS/e-commerce platform development, and embedded technical talent / staff augmentation to move AI from experimentation into production. With innovation hubs in Chicago and Karachi, Ontrac focuses on delivering production-ready AI and cloud solutions that drive measurable business outcomes.

📋 Description

• Be on an on-call rotation responding to production availability incidents and support service engineers with customer incidents • Use on-call shifts to prevent incidents from recurring • Run infrastructure with Ansible, Puppet, Terraform, and Kubernetes • Configure monitoring and alerting to report symptoms rather than outages • Document every action so findings become repeatable actions and automation • Improve the deployment process • Design, build, and maintain core infrastructure scaling to hundreds of thousands of concurrent users • Debug production issues across services and stack levels • Plan infrastructure growth • Code infrastructure automation with Ansible and Terraform • Improve Prometheus monitoring or build new metrics • Help release managers deploy and fix new application software versions • Plan and execute migration from AWS virtual machines to cloud-native, container-based deployments on Kubernetes (EKS) • Develop relationships with product groups and define SRE KPIs

🎯 Requirements

• Think cloud-first regardless of public-cloud provider • Think security-first • Understand systems, including edge cases, failure modes, behaviours, and specific implementations • Knowledge of Linux and Windows • Knowledge of configuration-management systems such as Ansible or Puppet • Strong programming skills in Python, Java, Golang, or Node.js • Ability to collaborate and communicate asynchronously and document work thoroughly • Go-for-it attitude and willingness to fix broken systems • Experience with Nginx, HAProxy, Docker, Kubernetes, Terraform, or similar technologies

Apply Now

Similar Jobs

🕒 April 21

Smart Working

51 - 200

💼 Consulting

🏥 Healthcare

📣 Marketing

Senior AWS DevOps Engineer responsible for building and scaling cloud-native infrastructure on AWS. Join a high-impact platform team in a fully remote role.

Ansible

AWS

Cloud

EC2

Grafana

Kubernetes

Linux

Packer

Postgres

Prometheus

Python

Splunk

Terraform

🕒 April 21

Flatgigs

1 - 10

💼 Consulting

🎯 Recruiter

👥 HR Tech

Senior DevOps Engineer focusing on MLOps at AHOY, building secure multi-cloud infrastructure and optimizing CI/CD pipelines.

Azure

Cloud

Google Cloud Platform

Kubernetes

Terraform

🕒 December 17, 2025

HR Ways - Hiring Tech Talent

11 - 50

💼 Consulting

📦 Logistics

👥 HR Tech

DevOps Engineer responsible for managing infrastructure & cloud operations for Fleet Management system. Work includes maintaining uptime and optimizing CI/CD pipelines for production systems.

AWS

Azure

Cloud

Docker

Google Cloud Platform

Jenkins

Kubernetes

Prometheus

Python

Terraform

🕒 September 29, 2025

Creative Chaos

201 - 500

💼 Consulting

📣 Marketing

📦 Logistics

Senior DevOps Engineer building and maintaining AWS cloud infrastructure and automation. Responsible for CI/CD, security, Linux administration, and environment deployments.

Ansible

AWS

Chef

Cloud

Docker

Jenkins

Kubernetes

Linux

Puppet

Python

Ruby

Subversion

Terraform

🕒 September 29, 2025

Creative Chaos

201 - 500

💼 Consulting

📣 Marketing

📦 Logistics

Senior DevOps Engineer responsible for Puppet automation, CI/CD, and cloud infrastructure operations. Ensuring secure, scalable environments and collaborating with developers for deployments.

Ansible

Chef

Cloud

Kubernetes

Linux

Puppet

Python

Ruby

Terraform