Senior Site Reliability Engineer

🔥 1 minute ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of PayNearMe

PayNearMe

201 - 500 employees

Founded 2009

💳 Fintech

☁️ SaaS

🤝 B2B

🔥 Funding within the last year

💰 $50M Series E - PayNearMe on 2025-09

Fintech • SaaS • B2B

PayNearMe is a payments technology company that provides a modern, end-to-end platform for businesses to accept, disburse and manage payments. The platform emphasizes Payment Experience Management, offering features such as modern payment processing, automation and self-service, exception management, and a cash-at-retail network. PayNearMe targets business customers across industries (auto & consumer lending, tolling, iGaming, buy-here-pay-here dealers, credit unions, mortgage servicing, law firms, etc. ) and advertises metrics including 16+ years in operation, $50B+ processed annually, 16,000+ businesses on its platform, and 62,000+ cash-at-retail locations. Money transmission services are provided by its subsidiaries PayNearMe MT, Inc. and PayNearMe Financial, Inc.

📋 Description

• Infrastructure Management: Design, implement, and maintain scalable and resilient infrastructure using Terraform for infrastructure as code, ensuring high availability and performance • Kubernetes and Containers: Deploy, manage, and optimize Kubernetes clusters and containerized applications using Docker. Implement best practices for container orchestration and management • Systems and Application Monitoring/Observability: Develop and maintain comprehensive monitoring and observability solutions using Datadog. Ensure detailed visibility into system performance and application health • SLOs and SLA Management: Define, monitor, and maintain Service Level Objectives (SLOs) and Service Level Agreements (SLAs) to ensure reliable and consistent service delivery • Incident Response and Troubleshooting: Respond to incidents, perform root cause analysis, and implement solutions to prevent recurrence. Participate in post-incident reviews and contribute to blameless postmortems • Reliability and Production Environment Management: Ensure the reliability and stability of our production environments. Continuously assess and improve system reliability, identifying and addressing potential points of failure • Automation and Scripting: Develop automation scripts and tools to reduce manual intervention and improve system reliability using Python, Bash, or Go. Implement and improve CI/CD pipelines • CI/CD Pipeline Management: Enhance and maintain continuous integration and continuous deployment pipelines using GitLab CI. Ensure seamless and reliable deployment processes • Capacity Planning and Scaling: Assist in capacity planning and ensure that systems are scalable to meet future demands. Implement auto-scaling strategies where applicable • Security and Compliance: Implement security best practices and ensure compliance with industry standards. Regularly review and update security policies and procedures • Collaboration and Support: Work closely with development teams to ensure reliability and scalability of new features and services. Provide technical support and guidance on infrastructure-related issues • Software Engineering for Operations: Develop and maintain internal tools and services that enhance the efficiency and reliability of our operations • On-Call Rotation: Participate in an on-call rotation to address production issues and collaborate in incident response efforts

🎯 Requirements

• +3 years of experience in SRE, DevOps, or a related role • Cloud Platform Experience: Proficient with cloud platforms such as AWS, GCP, or Azure Experience with EC2, RDS, VPCs, and security groups is essential. • Kubernetes and Containers: Strong experience with Kubernetes and Docker, including deployment, scaling, and management of containerized applications • Infrastructure as Code: Expert in using Terraform for infrastructure as code. Proficient with configuration management tools such as Ansible, Puppet, or Chef • Monitoring and Observability: Extensive experience with monitoring and observability tools like Datadog, Prometheus, Grafana, ELK stack, or Splunk. Skilled in setting up detailed monitoring and logging systems • SLOs and SLA Management: Proven ability to define, monitor, and maintain SLOs and SLAs to ensure reliable service delivery • Scripting and Automation: Strong skills in scripting languages like Python, Bash, or Go. Experience automating repetitive tasks and processes • CI/CD Practices: Familiarity with GitLab CI or similar tool for continuous integration and deployment. Experience in setting up and managing pipelines • Production Environments: Experience supporting production environments running Go or Ruby/Rails applications • Tool Development: Ability to write and update tools to support infrastructure and application management, demonstrating the principle that “SRE is what happens when you ask a software engineer to design an operations team • DevOps Best Practices: Deep understanding of DevOps principles, practices, and tools to drive continuous improvement in the software development lifecycle • Soft Skills: Strong organizational skills, attention to detail, and the ability to work collaboratively in a team environment. Excellent documentation skills to ensure accurate and detailed records • Problem-Solving Ability: Excellent analytical and problem-solving skills to diagnose and resolve complex system issues quickly and effectively.

🏖️ Benefits

• Competitive salary and benefits with growth-company options grant • Fast- paced and professional work culture • Stock options with standard startup vesting - 1 year cliff; 4 years total • $50 monthly communication expense stipend to go towards your phone/internet bill • $250 stipend to enhance your WFH setup • Reimbursement for peripheral equipment: monitor (up to $400), keyboard and mouse (up to $200) • Premium medical benefits including vision and dental (100% coverage for employees) • Company-sponsored life and disability insurance • Paid parental bonding leave • Paid sick leave, jury duty, bereavement • 401k plan • Flexible Time Off (our team members typically take off ~3-4 weeks per year) • Volunteer Time Off • 13 scheduled holidays

Apply Now

Similar Jobs

🔥 24 minutes ago

Made4net

51 - 200

📦 Logistics

☁️ SaaS

🏢 Enterprise

Cloud Operations Engineer supporting AWS infrastructure for supply chain software solutions. Monitoring systems and ensuring reliability in a global operations team.

Ansible

AWS

Cloud

DNS

EC2

Grafana

ITSM

Linux

Microservices

Oracle

Postgres

Python

Terraform

🔥 52 minutes ago

Mercadona

10,000+ employees

🛒 Retail

🍽️ Food & Beverage

DevOps Prime overseeing GCH’s cloud infrastructure while directing vendor resources and ensuring security. Responsible for CI/CD strategies, observability, and incident response across platforms.

AWS

Cloud

Terraform

🔥 1 hour ago

Sycurio

51 - 200

☁️ SaaS

🔐 Security

📋 Compliance

Deployment Engineer for Sycurio solutions deployment and testing. Involves supporting installations and training customer support engineers.

AWS

DNS

EC2

Firewalls

JavaScript

Linux

SOAP

Splunk

TCP/IP

VMware

VoIP

🔥 1 hour ago

Otoe Missouria Group

11 - 50

🏛️ Government

🔒 Cybersecurity

💼 Consulting

DevSecOps Engineer supporting federal agency application modernization efforts in Washington, DC. Building secure, efficient pipelines for mission-critical systems.

Azure

Cloud

Django

Flask

Python

🔥 1 hour ago

Imagineeer

11 - 50

🏛️ Government

🔒 Cybersecurity

💼 Consulting

DevOps Engineer specializing in SharePoint and Microsoft 365 platform deployment automation. Collaborating with teams to enhance cybersecurity, IT modernization in a federal context.

Azure

Cyber Security

Terraform