Site Reliability Engineer (SRE) – UI/UX

🔥 0 minutes ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Software Mind

Software Mind

1001 - 5000 employees

Founded 1999

🤖 Artificial Intelligence

☁️ SaaS

📡 Telecommunications

💰 Private Equity Round on 2020-12

Artificial Intelligence • SaaS • Telecommunications

Software Mind is a technology company that specializes in software development and digital transformation services. With a focus on AI and cloud solutions, the company offers a wide range of services including custom software development, mobile app development, and cloud consulting. Software Mind serves various industries such as financial services, telecom, biotech, and media, providing tailored solutions to accelerate digital transformations and business growth globally.

📋 Description

• Support the deployment, operations, and ongoing maintenance of production services running on Kubernetes • Monitor service health, availability, and performance • Investigate and troubleshoot production incidents using logs, monitoring, and debugging tools • Perform log analysis and incident debugging using Splunk • Identify service issues and collaborate with engineering teams to support timely resolution • Participate in incident response and production support activities • Perform first-level debugging of UI-related issues involving Web Components • Support service reliability and continuous improvement initiatives • Assist with CI/CD pipelines and cloud-native application operations when needed • Work effectively within a client-directed backlog and established priorities

🎯 Requirements

• 4+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, Production Support, or a related role • Hands-on experience supporting deployment, operations, and ongoing maintenance of production services running on Kubernetes • Experience monitoring service health, troubleshooting production issues, and supporting service reliability • Proficiency with Splunk for log analysis and incident debugging • Experience participating in production incident response and root-cause analysis • Working knowledge of Web Components and ability to perform first-level debugging of UI-related issues • Strong troubleshooting, analytical, and problem-solving skills • Experience collaborating with software engineering and cross-functional teams • Ability to work independently and effectively within a client-directed backlog • Excellent written and spoken English, at least B2 level • Preferred: experience supporting CI/CD pipelines • Preferred: familiarity with multi-tenant services • Preferred: experience with cloud-native application operations • Preferred: experience supporting high-availability enterprise or SaaS platforms • Preferred: familiarity with additional monitoring and observability tools • Preferred: experience with cloud platforms such as AWS, Azure, or GCP • Preferred: familiarity with Docker and Helm

🏖️ Benefits

• Competitive salary • Laptop • Professional development and training opportunities • Work with cutting-edge cloud and container technologies • Flexible work arrangements • Collaborative team environment • Impact on organization-wide digital transformation initiatives

Apply Now

Similar Jobs

🔥 1 hour ago

Autodesk

10,000+ employees

🏗️ Construction

🏭 Manufacturing

💼 Consulting

Senior DevOps Developer building reliable AWS, Kubernetes, and MongoDB services for Autodesk Construction Solutions. Improving automation, observability, security, disaster recovery, and production reliability for construction software customers.

🔥 6 hours ago

High Tech Genesis

51 - 200

📦 Logistics

📣 Marketing

🏭 Manufacturing

Cloud DevOps Engineer building AWS infrastructure, CI/CD pipelines, and automation for High Tech Genesis. Supporting containers, event-driven systems, service mesh, and observability.

🕒 Yesterday

S&P Global

10,000+ employees

💼 Consulting

📦 Logistics

📣 Marketing

DevOps Engineer building AWS infrastructure and automated systems for S&P Global’s financial data and technology solutions. Supporting resilient applications through Terraform, CI/CD, containerization, monitoring, and cloud operations.

🕒 4 days ago

Yelp

1001 - 5000

🍽️ Food & Beverage

🏨 Hospitality

📣 Marketing

Site Reliability Engineer operating Yelp’s Kafka and Flink streaming platform across Canada. Automating cluster management, scaling, upgrades, migrations, and incident recovery for real-time data systems.

🕒 5 days ago

Yelp

1001 - 5000

🍽️ Food & Beverage

🏨 Hospitality

📣 Marketing

Site Reliability Engineer operating Yelp’s Kafka and Flink streaming infrastructure across Canada. Automating cluster operations, upgrades, scaling, and incident recovery for real-time data systems.