Senior Site Reliability Operations Engineer – Finance

🔥 1 hour ago

🇩🇴 Dominican Republic – Remote

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

👻 Ghost score 25%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Truelogic Software

Truelogic Software

501 - 1000 employees

Founded 2004

☁️ SaaS

🤝 B2B

🏢 Enterprise

SaaS • B2B • Enterprise

Truelogic Software is a nearshore software development company specializing in agile staff augmentation services. They focus on providing custom outsourced software development with a team of highly skilled engineers from Latin America. Truelogic Software partners with both startups and Fortune 500 companies, offering solutions that align with their clients' time zones and ensuring high-quality outcomes through collaboration and responsiveness. With a presence in over 25 countries, Truelogic emphasizes remote work for better quality of life, and their engineers are experienced in various industries, delivering a wide range of successful projects globally.

📋 Description

• Lead incident response as Incident Commander, coordinating teams, communications, and service restoration • Produce executive-level incident reports and run root-cause analyses • Drive continuous improvement • Monitor and improve observability using AWS CloudWatch, New Relic, Nagios, and SumoLogic • Reduce alert noise and observability gaps • Provide hands-on system support across Linux and Windows environments • Troubleshoot complex infrastructure issues • Manage and execute deployments via Jenkins, GitLab, or similar CI/CD platforms • Own infrastructure initiatives including migrations, upgrades, and process improvements • Enforce change management and risk assessment for production changes • Maintain documentation and standard operating procedures • Liaise between engineering teams and external vendors • Participate in a one-week on-call rotation, including potential critical incident call-ins between 6:00 PM and 6:00 AM PT

🎯 Requirements

• 5+ years of experience in Windows and Linux environments with proven troubleshooting capabilities • Strong knowledge of monitoring tools like AWS CloudWatch, New Relic, Nagios, SumoLogic • Practical experience with CI/CD tools such as Jenkins and GitLab • Practical experience with backup tools such as CommVault and AWS Backup • Strong scripting skills in PowerShell, Python, or equivalent • Outstanding communication skills, especially under pressure, including executive reporting • Experience in high-paced environments and with on-call support models • Autonomous and proactive attitude; capable of managing complex tasks independently • Reliable internet connection and a laptop • Currently based in Latin America

🏖️ Benefits

• 100% Remote Work • Highly Competitive USD Pay • Paid Time Off • Work with Autonomy • Work with Top American Companies • Engagement activities • Work-life balance support • Collaboration with a diverse, multicultural global network • Opportunity to work with senior, seasoned professionals

Apply Now