Senior SRE – Production Reliability Engineer

Job not on LinkedIn

🔥 15 minutes ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Symphony Solutions

Symphony Solutions

501 - 1000 employees

Founded 2008

🤖 Artificial Intelligence

🎮 Gaming

☁️ SaaS

Artificial Intelligence • Gaming • SaaS

Symphony Solutions is a software company that builds AI-powered platforms and services focused on iGaming and enterprise software. Their public site highlights products such as BetSymphony (a turnkey iGaming / sportsbook and online casino platform), BetHarmony (an AI-powered conversational agent and bet recommendation system), Harmony (an AI assistant for customer service), and Symphony Cloud (a cloud development environment and DevOps tooling). They market AI-driven business transformation services and platform solutions for operators and enterprises, and also reference a corporate charity initiative.

📋 Description

• **Responsibilities** • - Own production reliability practices together with DevOps and engineering teams. • - Monitor production health and improve observability. • - Support releases, hotfixes, rollback, and post-release watch. • - Validate deployment readiness and rollback readiness. • - Participate in incident response and RCA/postmortem. • - Maintain runbooks, dashboards, alerting rules, and operational documentation. • - Ensure release evidence is traceable and audit-ready. • - Coordinate with Dev, QA, Product, Support, Release Manager, and DevOps during production events.

🎯 Requirements

• **Must Have: Technical Skills** • - **Linux** — confident troubleshooting in terminal, logs, processes, networking basics, resource usage. • - **Kubernetes / GKE** — strong hands-on understanding of workloads, pods, services, ingress/gateway, probes, RBAC, resources, autoscaling, and troubleshooting. • - **GCP** — practical experience with cloud infrastructure, especially around GKE, IAM, networking, load balancing, artifact registry, and production diagnostics. • - **Docker / Containers** — images, registries, runtime debugging, container lifecycle. • - **Helm** — understands Helm releases, values, deployment state, and rollback. • - **GitOps / FluxCD** — able to understand and troubleshoot GitOps deployment flow, drift, image automation, and Git-based rollback. • - **CI/CD understanding** — can investigate failed pipelines and deployment issues; does not need to be the main pipeline builder if DevOps owns that. • - **Observability** — Prometheus, Alertmanager, Grafana; understands metrics, logs, traces, dashboards, alerting, and production monitoring. • - **Networking fundamentals** — DNS, load balancers, ingress, gateways, TLS, routing, firewall/security rules. • - **SLI / SLO / SLA** — can define and apply reliability targets, not just explain the terms. • - **Incident response** — experience with production incidents, rollback, service recovery, RCA/postmortem. • - **Databases / messaging operational basics** — PostgreSQL, Couchbase, Kafka, Elasticsearch/ELK or similar; enough to monitor health, migrations, indexes, topics, and failure signals. • - **Security / compliance basics** — secrets, IAM/RBAC, audit trail, release/change evidence. • **Must Have: Reliability / Operational Skills** • - Can assess if a release is technically safe to proceed. • - Can define what should be monitored during and after deployment. • - Can verify production health after release. • - Can prepare or validate rollback plans. • - Can lead or support post-release watch. • - Can improve runbooks and incident response procedures. • - Can identify gaps in dashboards, alerts, deployment checks, and release evidence. • - Can work with DevOps without duplicating their ownership. • **Must Have: Soft Skills** • - **Ownership and accountability** — follows production issues through to closure. • - **Calm under pressure** — can operate clearly during incidents, failed deployments, and rollbacks. • - **Clear communication** — gives concise updates: impact, current action, ETA, risk, decision needed. • - **Structured troubleshooting** — breaks issues down by app, infra, network, database, config, release diff, external dependency. • - **Risk awareness** — can say when a release is risky and explain why. • - **Collaboration** — works well with Dev, QA, Product, Support, Release Manager, and DevOps. • - **Documentation discipline** — maintains runbooks, rollback steps, incident timelines, and operational evidence. • - **Blameless mindset** — focuses on root cause and prevention. • - **Proactivity** — finds missing alerts, dashboards, runbooks, and process gaps before incidents. • - **Prioritization** — separates real production impact from alert noise.

Apply Now