Middle SRE – Production Reliability Engineer

🔥 2 hours ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Symphony Solutions

Symphony Solutions

501 - 1000 employees

Founded 2008

🤖 Artificial Intelligence

🎮 Gaming

☁️ SaaS

Artificial Intelligence • Gaming • SaaS

Symphony Solutions is a software company that builds AI-powered platforms and services focused on iGaming and enterprise software. Their public site highlights products such as BetSymphony (a turnkey iGaming / sportsbook and online casino platform), BetHarmony (an AI-powered conversational agent and bet recommendation system), Harmony (an AI assistant for customer service), and Symphony Cloud (a cloud development environment and DevOps tooling). They market AI-driven business transformation services and platform solutions for operators and enterprises, and also reference a corporate charity initiative.

📋 Description

• Own production reliability practices together with DevOps and engineering teams. • Monitor production health and improve observability. • Support releases, hotfixes, rollback, and post-release watch. • Validate deployment readiness and rollback readiness. • Participate in incident response and RCA/postmortem. • Maintain runbooks, dashboards, alerting rules, and operational documentation. • Ensure release evidence is traceable and audit-ready. • Coordinate with Dev, QA, Product, Support, Release Manager, and DevOps during production events.

🎯 Requirements

• Linux — confident troubleshooting in terminal, logs, processes, networking basics, resource usage. • Kubernetes / GKE — strong hands-on understanding of workloads, pods, services, ingress/gateway, probes, RBAC, resources, autoscaling, and troubleshooting. • GCP — practical experience with cloud infrastructure, especially around GKE, IAM, networking, load balancing, artifact registry, and production diagnostics. • Docker / Containers — images, registries, runtime debugging, container lifecycle. • Helm — understands Helm releases, values, deployment state, and rollback. • GitOps / FluxCD — able to understand and troubleshoot GitOps deployment flow, drift, image automation, and Git-based rollback. • CI/CD understanding — can investigate failed pipelines and deployment issues; does not need to be the main pipeline builder if DevOps owns that. • Observability — Prometheus, Alertmanager, Grafana; understands metrics, logs, traces, dashboards, alerting, and production monitoring. • Networking fundamentals — DNS, load balancers, ingress, gateways, TLS, routing, firewall/security rules. • SLI / SLO / SLA — can define and apply reliability targets, not just explain the terms. • Incident response — experience with production incidents, rollback, service recovery, RCA/postmortem. • Databases / messaging operational basics — PostgreSQL, Couchbase, Kafka, Elasticsearch/ELK or similar; enough to monitor health, migrations, indexes, topics, and failure signals. • Security / compliance basics — secrets, IAM/RBAC, audit trail, release/change evidence. • Must Have: Reliability / Operational Skills • Can assess if a release is technically safe to proceed. • Can define what should be monitored during and after deployment. • Can verify production health after release. • Can prepare or validate rollback plans. • Can lead or support post-release watch. • Can improve runbooks and incident response procedures. • Can identify gaps in dashboards, alerts, deployment checks, and release evidence. • Can work with DevOps without duplicating their ownership. • Must Have: Soft Skills • Ownership and accountability — follows production issues through to closure. • Calm under pressure — can operate clearly during incidents, failed deployments, and rollbacks. • Clear communication — gives concise updates: impact, current action, ETA, risk, decision needed. • Structured troubleshooting — breaks issues down by app, infra, network, database, config, release diff, external dependency. • Risk awareness — can say when a release is risky and explain why. • Collaboration — works well with Dev, QA, Product, Support, Release Manager, and DevOps. • Documentation discipline — maintains runbooks, rollback steps, incident timelines, and operational evidence. • Blameless mindset — focuses on root cause and prevention. • Proactivity — finds missing alerts, dashboards, runbooks, and process gaps before incidents. • Prioritization — separates real production impact from alert noise. • Nice To Have • Terraform / IaC experience. • Advanced GCP infrastructure design. • ArgoCD or other GitOps tools. • Advanced database or performance tuning. • Service mesh experience. • On-call rotation experience with PagerDuty/Opsgenie or similar. • Experience facilitating RCA/postmortem sessions. • Betting/gaming domain experience. • Good written English for incident updates, release notes, and audit evidence. • Basic understanding of AI-related concepts: agents, skills, MCP.

Apply Now