Application Site Reliability Engineer, SRE

Job not on LinkedIn

🔥 0 minutes ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of CXM Direct

CXM Direct

51 - 200 employees

Founded 2015

💸 Finance

💳 Fintech

Finance • Fintech

CXM Direct is a global forex and CFD broker offering a wide range of trading instruments including currencies, commodities, indices, energy products, and cryptocurrencies. With advanced trading platforms such as MetaTrader 4 and 5, CXM Direct provides innovative trading solutions tailored to the needs of professional traders. The company offers various account types and services such as leverage policies, instant deposits, and withdrawals, ensuring a seamless trading experience. Regulated in multiple jurisdictions, CXM Direct emphasizes safety and sophistication in trading, positioning itself as a pioneer in the online trading environment.

📋 Description

• Own the day-to-day reliability of .NET/C# services running on Windows. • Participate in the on-call rotation for production trading systems and lead incident response during service disruptions. • Investigate production incidents, perform root cause analysis, and implement preventive actions to eliminate recurring issues. • Build and maintain Grafana dashboards, Prometheus alerts, and operational health views across applications, infrastructure, and databases. • Instrument .NET services to improve telemetry, metrics, logging, and visibility into service health and customer impact. • Define, implement, and monitor Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets. • Troubleshoot issues across .NET/C# applications, Windows Server, Aurora PostgreSQL databases, AWS infrastructure, CI/CD pipelines and deployments. • Improve deployment safety, release automation, and rollback strategies. • Partner with developers to improve application operability, resilience, and fault isolation. • Automate operational tasks through scripting and infrastructure automation. • Create and maintain runbooks, operational documentation, and incident response procedures. • Continuously improve monitoring, alert quality, automation, and platform reliability.

🎯 Requirements

• 3–5 years of experience in .NET/C# applications in production • Strong experience debugging and supporting .NET/C# applications in production. • Hands-on experience with Windows Server environments. • Strong PowerShell scripting skills. • Experience with Python or Bash. • Experience with Grafana, Prometheus, and Loki (or equivalent monitoring and observability tools). • Solid understanding of metrics, logging, tracing, and alerting best practices. • Experience with modern CI/CD pipelines. • Knowledge of deployment strategies, release automation, and rollback mechanisms. • Experience working with AWS. • Hands-on experience with Terraform or other Infrastructure as Code (IaC) tools. • Experience troubleshooting and supporting Aurora PostgreSQL or other relational database platforms. • Practical experience with SLIs & SLOs, Error Budgets, Incident Response, Root Cause Analysis (RCA), Alert Design, Production Operations.

🏖️ Benefits

• Work on mission-critical trading infrastructure that directly impacts customers. • Solve challenging reliability and scalability problems in a real-time environment. • Build world-class observability, automation, and deployment practices. • Collaborate with experienced engineers in a modern engineering culture. • Influence reliability strategy and engineering best practices across the platform.

Apply Now

Similar Jobs

🕒 July 6

Senior Application Security Engineer joining Canary Technologies to embed security into software development lifecycle. Collaborate with engineering teams for secure practices and tooling.

AWS

Cloud

JavaScript

Kubernetes

Python

SDLC

Terraform

Go

🕒 June 23

Workana

51 - 200

👥 HR Tech

🏪 Marketplace

🎯 Recruiter

AI Application Engineer developing AI-powered applications and workflows for medxprts.ai. Collaborating with engineering teams to build features and improve operational workflows.

AWS

Cloud

Docker

Google Cloud Platform

Kubernetes

Python