Site Reliability Engineer

🔥 0 minutes ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Ontrac Solutions

Ontrac Solutions

11 - 50 employees

Founded 2010

🤖 Artificial Intelligence

💼 Consulting

🤝 B2B

Artificial Intelligence • Consulting • B2B

Ontrac Solutions is an AI-first technology services firm that helps organizations design, implement, and scale AI-driven systems, cloud architectures, and enterprise platform integrations. They provide AI strategy and generative AI implementation (including RAG and intelligent search), cloud migration and optimization (multi-cloud architecture, Kubernetes, Terraform, GitOps, and FinOps), data and integration engineering, CRM/CMS/e-commerce platform development, and embedded technical talent / staff augmentation to move AI from experimentation into production. With innovation hubs in Chicago and Karachi, Ontrac focuses on delivering production-ready AI and cloud solutions that drive measurable business outcomes.

📋 Description

• Be on an on-call rotation responding to production availability incidents and support service engineers with customer incidents • Use on-call shifts to prevent incidents from recurring • Run infrastructure with Ansible, Puppet, Terraform, and Kubernetes • Configure monitoring and alerting to detect symptoms rather than outages • Document every action so findings become repeatable actions and automation • Improve the deployment process • Design, build, and maintain core infrastructure scaling to hundreds of thousands of concurrent users • Debug production issues across services and stack levels • Plan infrastructure growth • Code infrastructure automation with Ansible and Terraform • Improve Prometheus monitoring or build new metrics • Help release managers deploy and fix new application software versions • Plan and execute migration from AWS virtual machines to cloud-native, container-based deployments on Kubernetes (EKS) • Develop relationships with product groups and define their SRE KPIs

🎯 Requirements

• Think cloud-first regardless of public cloud provider • Think security-first • Understand systems, including edge cases, failure modes, behaviors, and specific implementations • Know Linux and Windows • Know configuration-management systems such as Ansible or Puppet • Strong programming skills in Python, Java, Golang, or Node.js • Collaborate and communicate asynchronously • Document work thoroughly • Have a go-for-it attitude and fix broken systems • Experience with Nginx, HAProxy, Docker, Kubernetes, Terraform, or similar technologies

Apply Now

Similar Jobs

🔥 7 hours ago

Akamai Technologies

5001 - 10000

🔒 Cybersecurity

Senior Site Reliability Engineer automating Linux infrastructure and improving reliability for Akamai’s distributed cloud and edge platform. Troubleshooting large-scale systems, deploying software safely, and enhancing monitoring and remediation.

Ansible

AWS

Azure

Cloud

Distributed Systems

Google Cloud Platform

Grafana

Linux

Prometheus

Python

SaltStack

Splunk

Terraform

Go

🕒 3 days ago

Ford Motor Company

10,000+ employees

📦 Logistics

💼 Consulting

📣 Marketing

Lead Site Reliability Engineer improving Ford’s automotive Marketing and Sales Tech platform from India. Enhancing observability, automation, resilience, and incident response.

Cloud

Google Cloud Platform

Java

JavaScript

Node.js

OpenShift

Python

Terraform

Go

🕒 3 days ago

Weekday

501 - 1000

👗 Fashion

🛒 Retail

🛍️ eCommerce

DevOps Engineer building scalable cloud infrastructure and CI/CD systems for a client. Managing containers, automation, monitoring, and reliability using AWS, Docker, and Kubernetes.

Ansible

AWS

Azure

Cloud

Docker

Google Cloud Platform

Grafana

Jenkins

Kafka

Kubernetes

Linux

Microservices

Prometheus

Python

Terraform

Go

🕒 4 days ago

Weekday (YC W21)

11 - 50

💼 Consulting

👥 HR Tech

☁️ SaaS

DevOps Engineer building cloud infrastructure, CI/CD pipelines, and containerized systems for a Weekday client. Automating provisioning, monitoring reliability, and deployment workflows using AWS, Docker, Kubernetes, and IaC tools.

Ansible

AWS

Azure

Cloud

Docker

Google Cloud Platform

Grafana

Jenkins

Kafka

Kubernetes

Linux

Microservices

Prometheus

Python

Terraform

Go

🕒 4 days ago

Empower

10,000+ employees

💸 Finance

💳 Fintech

👥 B2C

Senior DevOps Engineer automating AWS infrastructure and CI/CD for Empower’s financial SaaS products. Improving reliability, security, and developer productivity through tooling, cloud services, and automation.

Angular

Ansible

Apache

AWS

Azure

Chef

Cloud

DNS

Docker

DynamoDB

EC2

Google Cloud Platform

Java

JavaScript

Jenkins

jQuery

Kubernetes

Linux

Maven

MySQL

NGINX

Node.js

Python

React

Ruby

Splunk

Terraform

Go