Site Reliability Engineer

🔥 8 minutes ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of CXM

CXM

201 - 500 employees

Founded 2015

💸 Finance

💳 Fintech

Finance • Fintech

CXM is an online financial services firm that operates as a brokerage platform offering Forex and CFD trading to clients. The company enforces regional access restrictions for regulatory compliance and provides risk warnings about the high-risk nature of leveraged trading. CXM also offers customer support for access and compliance inquiries.

📋 Description

• Own the day-to-day reliability of our .NET/C# services running on Windows. • Participate in the on-call rotation for production trading systems and lead incident response during service disruptions. • Investigate production incidents, perform root cause analysis, and implement preventive actions to eliminate recurring issues. • Build and maintain Grafana dashboards, Prometheus alerts, and operational health views across applications, infrastructure, and databases. • Instrument .NET services to improve telemetry, metrics, logging, and visibility into service health and customer impact. • Define, implement, and monitor Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets. • Troubleshoot issues across .NET/C# applications, Windows Server, Aurora PostgreSQL databases, and AWS infrastructure. • Improve deployment safety, release automation, and rollback strategies. • Partner with developers to improve application operability, resilience, and fault isolation. • Automate operational tasks through scripting and infrastructure automation. • Create and maintain runbooks, operational documentation, and incident response procedures. • Continuously improve monitoring, alert quality, automation, and platform reliability.

🎯 Requirements

• 3–5 years of experience in Site Reliability Engineering or related field • Strong experience debugging and supporting .NET/C# applications in production. • Hands-on experience with Windows Server environments. • Strong PowerShell scripting skills. • Experience with Python or Bash. • Experience with Grafana, Prometheus, and Loki (or equivalent monitoring and observability tools). • Experience with modern CI/CD pipelines. • Knowledge of deployment strategies, release automation, and rollback mechanisms. • Experience working with AWS. • Hands-on experience with Terraform or other Infrastructure as Code (IaC) tools. • Experience troubleshooting and supporting Aurora PostgreSQL or other relational database platforms. • Practical experience with SLIs & SLOs, Error Budgets, Incident Response, Root Cause Analysis (RCA), and Alert Design.

🏖️ Benefits

• Work on mission-critical trading infrastructure that directly impacts customers. • Solve challenging reliability and scalability problems in a real-time environment. • Build world-class observability, automation, and deployment practices. • Collaborate with experienced engineers in a modern engineering culture. • Influence reliability strategy and engineering best practices across the platform.

Apply Now

Similar Jobs

🔥 3 hours ago

NVIDIA

10,000+ employees

🏥 Healthcare

🏭 Manufacturing

🤖 Artificial Intelligence

DevOps Engineer supporting NVIDIA’s Rapids project for AI and data science initiatives. Collaborating with teams to ensure high-quality software releases and infrastructure maintenance.

🔥 10 hours ago

Softgic

51 - 200

💼 Consulting

🔒 Cybersecurity

DevOps Specialist managing Google Cloud infrastructure at Softgic. Automating and optimizing cloud systems and collaborating with development teams.

🗣️🇪🇸 Spanish Required

🔥 13 hours ago

Global Enterprise Services, LLC (GES)

11 - 50

💼 Consulting

📦 Logistics

Reliability Engineer responsible for cloud platform performance and incident response, managing compliance. Requires strong technical expertise and 8 years of experience.

🔥 19 hours ago

IPolarity

51 - 200

💼 Consulting

🏥 Healthcare

📦 Logistics

Senior Site Reliability Engineer managing AWS GovCloud and multi-account environments. Building automation and overseeing cloud governance in a fully remote setup.

🔥 20 hours ago

Mirantis

501 - 1000

💼 Consulting

🏥 Healthcare

📦 Logistics

Senior DevOps Engineer handling high-performance storage for AI platforms at Mirantis. Integrating and operating storage solutions within Kubernetes environments.