Staff Software Engineer – Reliability & Platform

🔥 0 minutes ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of One Identity

One Identity

501 - 1000 employees

🔒 Cybersecurity

☁️ SaaS

🏢 Enterprise

💰 $3M Venture Round on 2004-07

Cybersecurity • SaaS • Enterprise

One Identity is a company that specializes in identity and access management solutions, focusing on protecting digital identities and simplifying user access across organizations. They offer a comprehensive suite of products designed to secure privileged access, govern user identities, and streamline compliance with regulations through automation. Their platform integrates AI-driven insights to enhance security and operational efficiency, supporting both on-premises and cloud environments.

📋 Description

• Own production operability by debugging complex issues, improving system visibility, and eliminating recurring problems at the source • Own production health for services — from detection through resolution to prevention • Improve mean time to detect (MTTD), mean time to resolve (MTTR), and recurrence rates for issues • Identify systemic issues and eliminate recurring problems through code fixes, architecture improvements, and better operational tooling • Improve observability across services — logs, metrics, and alerting — for faster diagnosis and resolution • Design and improve debugging workflows, runbooks, and internal tooling for engineers • Reduce operational burden by making systems easier to understand, operate, and troubleshoot • Partner closely with product teams to feed production learnings back into design and development • Reduce support and incident load by addressing root causes and improving system design, not just resolving individual issues

🎯 Requirements

• 4+ years of software engineering experience with ownership of production systems, reliability, or operational improvements • Strong backend development experience (Ruby, Node.js, or similar) • Solid understanding of REST APIs, service contracts, and software design principles • Experience working across backend services, APIs, and production systems • Experience building and operating services in AWS or similar cloud environments. • Experience with observability, production debugging, and incident response. • Willingness to participate in a mandatory 24/7 on-call rotation. • Experience responding to production incidents and contributing to reliability improvements. • Experience using, or strong interest in, AI-powered development tools (e.g., GitHub Copilot, ChatGPT, Cursor).

🏖️ Benefits

• Competitive salary • Flexible working hours • Professional development budget • Home office setup allowance • Global team events

Apply Now

Similar Jobs

🕒 Yesterday

Ping Identity

1001 - 5000

🔒 Cybersecurity

☁️ SaaS

🏢 Enterprise

Staff Site Reliability Engineer involved in all aspects of Cloud services at Ping Identity. Design, deploy, and maintain infrastructure for a major identity platform.

🕒 June 29

Grafana Labs

501 - 1000

🏢 Enterprise

☁️ SaaS

🤖 Artificial Intelligence

Staff Software Engineer ensuring Grafana Cloud database reliability for high-value customers. Collaborate with product teams and lead incident response efforts for exceptional service quality.

🕒 June 19

Reddit, Inc.

501 - 1000

👥 B2C

📱 Media

🌍 Social Impact

Staff Site Reliability Engineer leading reliability initiatives across Ads domains at Reddit. Working to improve reliability, scalability, and operational efficiency in Reddit's advertising ecosystem.

🕒 June 11

Advanced Solutions International, Inc.

201 - 500

🤝 B2B

🤝 Non-profit

DevOps Reliability Engineer ensuring performance, scalability, and reliability of Azure-based SaaS platform at ASI. Collaborating with engineering teams to improve system efficiency and resilience.