Lead Application Reliability Engineer

Job not on LinkedIn

🔥 0 minutes ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of InnoData

InnoData

2 - 10 employees

Founded 2019

🤝 B2B

💼 Consulting

🌍 Social Impact

B2B • Consulting • Social Impact

InnoData is INNOvation DATA SCS, an Italian social cooperative based in Foggia that identifies itself as a provider of technological solutions. The company website (currently under maintenance) highlights "Soluzioni tecnologiche" (technological solutions) and emphasizes social impact ("Impatto sociale"). Contact details listed include Via Francesco Crispi 65, 71121 Foggia, Italy. Based on the available information, InnoData appears to operate at the intersection of technology and social impact, likely offering tech-focused services to other organizations.

📋 Description

• Provide production support and maintenance for enterprise applications hosted on Google App Engine • Triage, diagnose, restore service, and close user-impacting incidents within agreed SLAs/SLOs • Troubleshoot application errors, latency, performance degradation, service failures, configuration issues, quota/scaling limits, and integration failures • Design, develop, and deliver enhancements to existing applications • Support microservices architecture, APIs, authentication, inter-service communication, and failure/retry behavior • Own build, release, and deployment activities across development, test, pre-production, and production • Manage App Engine versions, traffic splitting/migration, canary and staged rollouts, rollbacks, configuration, and scaling • Perform root-cause analysis and implement sustainable fixes • Build and maintain monitoring, alerting, logging, dashboards, and error reporting using Cloud Monitoring, Cloud Logging, Error Reporting, and Cloud Trace • Support platform, framework, library, dependency, and runtime upgrades • Support IAM, service accounts, access controls, secrets management, and operational governance • Participate in change management, release-readiness reviews, and on-call/rotational support • Collaborate with client and cross-functional Application Engineering, Product, QA, Data Engineering, Infrastructure, and Platform teams • Create and maintain technical documentation, runbooks, troubleshooting guides, deployment procedures, and support playbooks

🎯 Requirements

• 3–7 years of experience in Application Support, Application Engineering, Software Engineering, Cloud Engineering, or a related role • Strong hands-on experience supporting production applications on Google Cloud Platform • Hands-on experience with Google App Engine deployment, configuration, scaling, versioning, and troubleshooting • Understanding of microservices architecture, REST/gRPC APIs, service-to-service communication and authentication, distributed tracing, and cross-service debugging • Experience with multi-environment deployments, release validation, rollback, and change control • Strong programming skills in one or more of Python, Java, Node.js/JavaScript, or Go • Working knowledge of SQL and application data stores including Cloud SQL, Firestore/Datastore, Cloud Spanner, or BigQuery • Understanding of GCP IAM, service accounts, permissions, monitoring, logging, alerting, and production operations • Experience with CI/CD pipelines and automated build/deployment tools such as Cloud Build, Jenkins, GitHub Actions, or GitLab CI • Experience troubleshooting complex production environments and performing root-cause analysis under time pressure • Ability to understand existing systems, codebases, services, configurations, and client-specific workflows • Strong communication and collaboration skills • Preferred: experience with large-scale enterprise applications, Cloud Run, GKE, Cloud Functions, Apigee/API Gateway, Pub/Sub, Cloud Tasks, Terraform, SRE practices, Docker, Kubernetes, frontend/full-stack applications, BI platforms, or AI/ML applications on GCP

Apply Now

Similar Jobs

🔥 2 hours ago

WinAir

51 - 200

📦 Logistics

💼 Consulting

🏭 Manufacturing

DevOps Specialist automating CI/CD and infrastructure for WinAir’s aviation maintenance software. Improving Jenkins, Ansible, Linux environments, deployments, and monitoring across development and production systems.

🔥 22 hours ago

Software Mind

1001 - 5000

🤖 Artificial Intelligence

☁️ SaaS

📡 Telecommunications

Senior SRE maintaining Kubernetes-based UI and AI service reliability for an enterprise cloud software company. Managing incidents, observability, deployments, and runtime troubleshooting in production.

🕒 Yesterday

Software Mind

1001 - 5000

🤖 Artificial Intelligence

☁️ SaaS

📡 Telecommunications

Senior SRE supporting Kubernetes production reliability for an enterprise cloud software company. Troubleshooting distributed services, observability, incidents, CI/CD, Node.js, and JVM/Java runtimes.

🕒 2 days ago

Software Mind

1001 - 5000

🤖 Artificial Intelligence

☁️ SaaS

📡 Telecommunications

Site Reliability Engineer operating Kubernetes UI services for a leading enterprise cloud software company. Monitoring reliability, troubleshooting incidents, and supporting production operations.

🇨🇦 Canada – Remote

💰 Private Equity Round on 2020-12

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🕒 2 days ago

Software Mind

1001 - 5000

🤖 Artificial Intelligence

☁️ SaaS

📡 Telecommunications

Site Reliability Engineer maintaining Kubernetes UI services for a leading enterprise software company. Monitoring production reliability, troubleshooting incidents, and supporting cloud-native operations.

🇨🇦 Canada – Remote

💰 Private Equity Round on 2020-12

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)