Site Reliability Engineer

🕒 July 21

🇨🇷 Costa Rica – Remote

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

👻 Ghost score 21%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of CSC Generation

CSC Generation

1001 - 5000 employees

🛒 Retail

🛍️ eCommerce

🤖 Artificial Intelligence

💰 Venture Round on 2019-01

Retail • eCommerce • Artificial Intelligence

CSC Generation is a retail technology platform that enhances revenue growth and unit margin management through automation and AI. The company manages a diverse portfolio of 325,000 products online and receives over 10 million monthly page views across its brands. CSC Generation specializes in retail, ecommerce, and wholesale, with a commitment to expanding by acquiring successful brands. Founded in 2016 by Justin Yoshimura, CSC Generation has acquired several well-known brands including One Kings Lane and Sur La Table, and continues to seek out new brands to join its network. The company offers extensive career opportunities and focuses on creating an inspiring and challenging work environment.

📋 Description

• Work on service resiliency, performance tuning, and system design across Backcountry's platform • Drive resolution of critical incidents and ensure fixes are methodically implemented through postmortems • Leverage AI-assisted engineering tools (Claude Code, GitHub Copilot, MCP-based agents) to investigate, automate, and ship fixes across infrastructure and application repositories • Reduce toil by designing and implementing automation • Partner with other Site Reliability Engineers, developers, and architects to evaluate and implement best practices for current and future workloads • Monitor system health and capacity, taking proactive action to fix problems before they occur • Collaborate with engineering teams to build, deploy, and support features • Build and maintain observability (metrics, logs, traces, profiles) and SLI/SLO instrumentation for Backcountry services • Participate in FinOps initiatives across GCP and AWS, including capacity planning and committed-use discount strategy • Participate in the on-call support rotation within the SRE team

🎯 Requirements

• 3+ years of experience supporting containerized production services, preferably running Kubernetes • 3+ years of experience with Infrastructure as Code (Terraform, AWS CDK, Ansible, etc.) • 3+ years of cloud experience operating in Google Cloud Platform and/or AWS (multi-cloud stack; Azure/Entra exposure is a plus) • Comfortable diagnosing issues and shipping bug fixes directly to application code (not just infrastructure) to keep services reliable and stable • Comfortable performing deep dives across both infrastructure and application/software git repositories to trace issues end-to-end • Proficient with AI-assisted coding tools (e.g., Claude Code, GitHub Copilot) and MCP-based agents, used to accelerate investigation, code review, and automation • Strong knowledge of scripting and programming languages (Bash, Python, and TypeScript/Node.js) • Experience managing Linux (any major distribution) in production environments • Excellent understanding of internet application protocols (DHCP, DNS, HTTPS, SSH, etc.) • Understanding of how DevOps (CI/CD) and SRE practices (SLOs, SLIs) apply to daily work • Hands-on experience with observability tooling (Grafana, Prometheus, Loki, OpenSearch, or equivalents) and SLI/SLO instrumentation • Experience with GitOps and Kubernetes packaging (ArgoCD, Helm, Kustomize) • Proactively track emerging technology trends and developments, evaluating which ones are worth bringing into engineering practice • Bachelor's degree in computer science or similar, or equivalent experience • Advanced-level English communication skills, both verbal and written.

🏖️ Benefits

• Competitive Benefits: We offer an attractive benefits package including primarily remote work, private medical and life insurance, additional paid time off, monthly allowances and reimbursements, employee discounts, and opportunities for professional growth.

Apply Now

Similar Jobs

🕒 July 9

Akamai Technologies

5001 - 10000

🔒 Cybersecurity

Site Reliability Engineer II creating solutions to improve automation and efficiency for systems. Optimizing workflows and collaborating on deployment and monitoring within Akamai's Compute products.

🇨🇷 Costa Rica – Remote

💵 ₡16.9M - ₡30.5M / year

💰 Post-IPO Equity on 2001-07

⏰ Full Time

🟢 Junior

🟡 Mid-level

⛑ DevOps & Site Reliability Engineer (SRE)

Ansible

Distributed Systems

Grafana

Kubernetes

Linux

Prometheus

Python

SaltStack

Terraform

Unix

Go

🕒 June 25

Akamai Technologies

5001 - 10000

🔒 Cybersecurity

Site Reliability Engineer optimizing large distributed content delivery systems for Akamai. Collaborating on reliability, scalability, and performance with product teams.

🇨🇷 Costa Rica – Remote

💵 ₡16.9M - ₡30.5M / year

💰 Post-IPO Equity on 2001-07

⏰ Full Time

🟢 Junior

🟡 Mid-level

⛑ DevOps & Site Reliability Engineer (SRE)

Ansible

Chef

Docker

Grafana

Kubernetes

Linux

Prometheus

Puppet

SaltStack

Terraform

Unix

🕒 May 30

GFT Technologies

10,000+ employees

💼 Consulting

🛡️ Insurance

🔒 Cybersecurity

DevOps & Network Engineer designing AWS-native networking capabilities for Cloud Networking initiative. Senior-level role requiring expertise in AWS networking and Terraform.

🗣️🇪🇸 Spanish Required

AWS

DNS

Terraform

🕒 May 4

Akamai Technologies

5001 - 10000

🔒 Cybersecurity

🏢 Enterprise

📱 Media

Maintenance and support of continuous integration pipelines at Akamai. Collaborating with product teams to streamline software development and application management.

🇨🇷 Costa Rica – Remote

💵 ₡15.4M - ₡32M / year

⏰ Full Time

🟢 Junior

🟡 Mid-level

⛑ DevOps & Site Reliability Engineer (SRE)

Ansible

Jenkins

Linux

Python

SaltStack

Terraform

Unix

🕒 April 29

GFT Technologies

10,000+ employees

💼 Consulting

🛡️ Insurance

🔒 Cybersecurity

🗣️🇪🇸 Spanish Required

Ansible

Jenkins

Linux

Postgres

Python

SQL

Terraform