Senior SRE, Managed Gateways

Job not on LinkedIn

🔥 5 minutes ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Kong Inc.

Kong Inc.

201 - 500 employees

Founded 2017

🔌 API

☁️ SaaS

🏢 Enterprise

💰 $100M Series D on 2021-02

API • SaaS • Enterprise

Kong Inc. is a company that provides a comprehensive API platform designed to facilitate API management, AI integration, and developer productivity. It offers solutions like Kong Gateway, Kong Konnect, and a variety of other tools targeted at managing and optimizing the API lifecycle. Kong's platform supports multi-cloud environments and is built to deliver high performance and security. It is notably recognized by Gartner as a leader in API management and supports innovations across industries like financial services, healthcare, and technology. The company emphasizes flexibility, security, and speed, making it a favored choice for enterprises looking to enhance their digital services through APIs. Kong also supports a robust community of developers and provides extensive integrations and plugins to streamline API management and operations.

📋 Description

• Lead, mentor, and inspire a high-performing team of Site Reliability Engineers dedicated to Kong's Managed Gateway offerings. • Architect and implement robust, scalable, and fault-tolerant cloud-native systems using technologies like Kubernetes, Golang, and major cloud providers. • Own the end-to-end operational lifecycle, from proactive monitoring and alerting to incident response and blameless post-mortems, ensuring continuous service improvement. • Drive a culture of developer delight by implementing automation, self-service tooling, and streamlined workflows for deploying and managing API gateways. • Define, track, and report on key SLOs and SLIs to ensure optimal performance and reliability of Managed Gateways. • Champion technical debt prevention and advocate for architectural best practices that enhance system resilience and reduce operational toil. • Collaborate cross-functionally with Product, engineering, and Customer Success to influence roadmap decisions and ensure operational readiness for new features. • Partner directly with enterprise customers — working alongside Product leadership, Professional Services, and Customer Success — to drive end-to-end onboarding and implementation of Cloud Gateways, and productize recurring implementation patterns into repeatable playbooks and platform capabilities. • Bring deep, cross-cloud breadth (AWS, GCP, Azure) to handle unique customer topologies and turn complex setups into successful, production-ready deployments.

🎯 Requirements

• Extensive experience as a Site Reliability Engineer, focusing on highly available and distributed systems. • Deep expertise with Kubernetes and cloud-native architectures, preferably across multiple public cloud providers (AWS, GCP, Azure). • Strong proficiency in Golang or similar modern programming languages for automation and tool development. • Proven track record in building and maintaining CI/CD pipelines and infrastructure as code (Terraform, Ansible). • In-depth knowledge of monitoring, logging, and alerting systems (e.g., Prometheus, Grafana, ELK stack, Datadog). • Experience with managed services, API gateways, or similar network infrastructure is highly desirable.

🏖️ Benefits

• health insurance • 401(k) plan • short and long term disability benefits • basic life and AD&D insurance

Apply Now

Similar Jobs

🔥 12 hours ago

Policy Reporter

51 - 200

📋 Compliance

DevOps Developer II at Policy Reporter improving CI/CD pipelines and collaborating with software teams. Hands-on development with Docker and Infrastructure-as-Code in a remote role.

🕒 3 days ago

Grafana Labs

501 - 1000

🏢 Enterprise

☁️ SaaS

🤖 Artificial Intelligence

Senior Software Engineer - SRE supporting Grafana Cloud customer databases for exceptional reliability. Collaborating with engineering teams on production systems and automation for high-SLA customers.

🕒 3 days ago

Semios

201 - 500

Senior Site Reliability Engineer ensuring the scalability and reliability of infrastructure for Semios Group. Leading initiatives in automation and team development for agricultural technology solutions.

🕒 6 days ago

3Pillar Global

1001 - 5000

☁️ SaaS

🏢 Enterprise

🤖 Artificial Intelligence

Senior DevOps Engineer responsible for operational excellence and innovative integration at 3Pillar. Driving advancements in cloud environments and deploying cutting-edge technologies while mentoring teams.

🕒 July 15

Autodesk

10,000+ employees

📱 Media

Senior DevOps Developer leading the evolution of a cloud-native platform for web applications and API integrations. Fostering security, compliance, and operational best practices with modern technologies.