Search Remote Jobs

Site Reliability Engineer

Job not on LinkedIn

🔥 15 minutes ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Sporttrade

Sporttrade

11 - 50 employees

Founded 2017

đź’Ľ Consulting

📣 Marketing

🎲 Gambling

đź’° $36M Funding Round on 2021-06

Consulting • Marketing • Gambling

Sporttrade is a sports betting marketplace that allows users to place, track, and sell their bets in real-time. Unlike traditional sportsbooks, Sporttrade uses a unique probability-based system rather than American odds, enabling bettors to easily understand the value of their bets. With multiple market makers competing for the best prices, Sporttrade offers a more refined betting experience through a modern interface designed for seamless betting. The platform aims to provide high limits and exceptional prices, catering to serious bettors over the age of 21.

đź“‹ Description

• Own the daily operation of the exchange trading lifecycle, including market startup and shutdown, enabling and disabling trading, pre- and post-session sanity checks, and capture of settlement, clearing, and trade-reporting artifacts • Participate in an on-call rotation for a live regulated marketplace • Lead incident response, drive incidents to resolution, author postmortems, and convert one-off fixes into runbooks and automation • Operate and improve the observability stack, including dashboards, alert quality, SLOs, and time-to-detection • Run and maintain hybrid infrastructure across Kubernetes clusters, cloud accounts, and geographically distributed on-premises datacenters • Automate infrastructure and operational procedures using Ansible, Terraform, and Jenkins pipelines, with secrets managed in HashiCorp Vault • Support the exchange data platform, including PostgreSQL, Kafka change-data-capture and streaming pipelines, Redis, backup/restore, and disaster recovery • Support market-maker and partner connectivity, conformance testing, and partner onboarding • Contribute to process improvement and establish policies and procedures for monitoring, incident management, change control, and exchange operations

🎯 Requirements

• 5+ years of experience in a Site Reliability Engineering, DevOps, production engineering, or technical operations role supporting a 24/7 production system • Strong Linux fundamentals and scripting ability • Experience supporting and debugging Java applications in production, including stack traces, thread dumps, JVM memory, garbage collection, logs, and metrics • Solid working knowledge of TCP/IP networking, including connections, ports, routing, and firewalls • Hands-on experience operating Kubernetes in production • Experience managing infrastructure as code with Terraform and Ansible • Experience with CI/CD pipelines using Jenkins or similar • Experience with observability tooling such as Datadog, Prometheus, and Grafana, or equivalents • Track record of being on-call for critical systems • Working knowledge of SQL and relational databases; PostgreSQL preferred • Self-starter able to deliver results with minimal guidance • Comfortable working independently and with a team • Excellent communication and organizational skills, especially written incident communication and documentation • Background or interest in trading, capital markets, exchange operations, or sports betting is a plus • Familiarity with exchange protocols is a plus • Previous experience in a regulated industry is a plus • Startup experience preferred but not required

🏖️ Benefits

• Medical, Dental, and Vision Benefits: Company pays 100% Employee premium and 50% Spouse & Dependent premiums • Short- & Long-Term Disability • Group Term Life and AD&D • Voluntary Life and AD&D • 401(k) Plan • Equity Options • Flexible time off • MacBooks issued to all employees

Apply Now

Similar Jobs

🔥 1 hour ago

NVIDIA

10,000+ employees

🏥 Healthcare

🏭 Manufacturing

🤖 Artificial Intelligence

AI Tools Engineer building LLM and AI/ML systems for NVIDIA’s global GeForce NOW service. Automating incident root-cause analysis and predicting operational trends from production data.

🔥 2 hours ago

SitusAMC

5001 - 10000

đź’Ľ Consulting

📦 Logistics

🏠 Real Estate

Site Reliability Engineer operating AWS cloud infrastructure for SitusAMC’s real estate technology solutions. Improving reliability, automation, observability, security, and application migrations.

🔥 4 hours ago

GE Vernova

10,000+ employees

đź’Ľ Consulting

📦 Logistics

🏭 Manufacturing

Senior Reliability Engineer improving embedded protection, control, and software products for GE Vernova’s decarbonization mission. Driving testing, field analytics, grid reliability, and cybersecurity compliance.

🔥 4 hours ago

SailPoint

1001 - 5000

đź’Ľ Consulting

🏥 Healthcare

📦 Logistics

Senior Staff DevOps Engineer scaling AWS Kubernetes infrastructure for SailPoint’s identity security platform. Leading enterprise service mesh adoption, PCI-compliant operations, and cloud-native reliability across global teams.

🔥 4 hours ago

Pacvue

501 - 1000

đź’Ľ Consulting

📣 Marketing

📦 Logistics

Senior DevOps Engineer building AWS/Kubernetes infrastructure and CI/CD systems for Pacvue’s commerce media platform. Improving reliability, security, observability, and developer productivity across engineering teams.