Site Reliability Engineer

🕒 July 26

🏈 North America – Remote

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

👻 Ghost score 34%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of CXM

CXM

201 - 500 employees

Founded 2015

💸 Finance

💳 Fintech

Finance • Fintech

CXM is an online financial services firm that operates as a brokerage platform offering Forex and CFD trading to clients. The company enforces regional access restrictions for regulatory compliance and provides risk warnings about the high-risk nature of leveraged trading. CXM also offers customer support for access and compliance inquiries.

📋 Description

• Own the day-to-day reliability of .NET/C# services running on Windows, starting with the in-house liquidity bridge connecting MetaTrader trading servers to external liquidity providers • Expand reliability impact across related trading and back-office services • Participate in the on-call rotation for production trading systems and lead incident response during service disruptions • Investigate production incidents, perform root cause analysis, and implement preventive actions • Build and maintain Grafana dashboards, Prometheus alerts, and operational health views across applications, infrastructure, and databases • Instrument .NET services to improve telemetry, metrics, logging, and visibility into service health and customer impact • Define, implement, and monitor SLIs, SLOs, and error budgets • Troubleshoot .NET/C# applications, Windows Server, Aurora PostgreSQL databases, AWS infrastructure, and CI/CD pipelines and deployments • Improve deployment safety, release automation, and rollback strategies • Partner with developers to improve application operability, resilience, and fault isolation • Automate operational tasks through scripting and infrastructure automation • Create and maintain runbooks, operational documentation, and incident response procedures • Continuously improve monitoring, alert quality, automation, and platform reliability

🎯 Requirements

• 3–5 years of experience (mid-level) • Strong experience debugging and supporting .NET/C# applications in production • Hands-on experience with Windows Server environments • Strong PowerShell scripting skills • Experience with Python or Bash • Experience with Grafana, Prometheus, and Loki (or equivalent monitoring and observability tools) • Solid understanding of metrics, logging, tracing, and alerting best practices • Experience with modern CI/CD pipelines • Knowledge of deployment strategies, release automation, and rollback mechanisms • Experience working with AWS • Hands-on experience with Terraform or other Infrastructure as Code (IaC) tools • Experience troubleshooting and supporting Aurora PostgreSQL or other relational database platforms • Practical experience with SLIs & SLOs, Error Budgets, Incident Response, Root Cause Analysis (RCA), Alert Design, and Production Operations • Preferred: Experience supporting high-availability or low-latency financial or trading systems • Preferred: Familiarity with MetaTrader environments or financial technology platforms • Preferred: Experience with distributed systems and microservices • Preferred: Knowledge of OpenTelemetry or similar observability frameworks • Preferred: Exposure to Docker, Kubernetes, or containerized environments

🏖️ Benefits

• Work on mission-critical trading infrastructure that directly impacts customers • Solve challenging reliability and scalability problems in a real-time environment • Build world-class observability, automation, and deployment practices • Collaborate with experienced engineers in a modern engineering culture • Influence reliability strategy and engineering best practices across the platform

Apply Now

Similar Jobs

🕒 May 14

Perry Street Software

51 - 200

💼 Consulting

🏥 Healthcare

✈️ Travel

Senior DevOps Engineer for Perry Street Software managing cloud infrastructure for LGBTQ+ dating apps. Collaborating with distributed teams to deliver reliable backend solutions.

🕒 April 1

Group 1001

501 - 1000

💼 Consulting

🏥 Healthcare

💸 Finance

DevOps Engineer optimizing CI/CD processes and maintaining AWS cloud infrastructure. Collaborative role focusing on automation, scalability, and cost optimization in cloud technologies.