
51 - 200 employees
💼 Consulting
📣 Marketing
📦 Logistics
Consulting • Marketing • Logistics
Tech Holding is a global technology solution company that provides Professional Services, Managed Services, and Staffing Solutions. Specializing in digital transformation and business growth, the company leverages its extensive industry knowledge and technological expertise to assist businesses in navigating the digital landscape. Tech Holding is committed to optimizing business operations and supplying the necessary talent to help clients achieve strategic objectives, making them a trusted partner for leading companies worldwide.
🔥 0 minutes ago
Improve your chances of getting an interview by checking your resume score before you apply.

51 - 200 employees
💼 Consulting
📣 Marketing
📦 Logistics
Consulting • Marketing • Logistics
Tech Holding is a global technology solution company that provides Professional Services, Managed Services, and Staffing Solutions. Specializing in digital transformation and business growth, the company leverages its extensive industry knowledge and technological expertise to assist businesses in navigating the digital landscape. Tech Holding is committed to optimizing business operations and supplying the necessary talent to help clients achieve strategic objectives, making them a trusted partner for leading companies worldwide.
• Establish performance, throughput, latency, and capacity baselines for critical customer and platform workflows • Define and maintain SLOs, error budgets, performance budgets, dashboards, alerts, and reliability thresholds • Instrument and analyze the full request path across application services, compute, storage, networking, databases, caches, queues, DNS, registry dependencies, and third-party services • Identify system bottlenecks and lead cross-functional remediation efforts with engineering teams • Build capacity models showing platform sustainability, emerging constraints, and additional scaling costs • Lead load, stress, soak, spike, failure, and recovery testing in representative environments • Develop realistic demand scenarios for major customers, partnerships, pilots, and high-volume events • Drive architecture hardening, graceful degradation, dependency-failure planning, and resilience improvements • Partner with Test Automation and Scalability Engineering on automated performance testing, regression coverage, and production release gates • Own technical readiness assessments for major pilots, partnerships, and production launches • Create operational runbooks for scale-up events, incidents, rollback, recovery, and dependency failures • Lead performance and reliability investigations during incidents and incorporate lessons into future engineering work • Make infrastructure cost, performance, and reliability tradeoffs visible to engineering and executive leadership • Recommend capacity and reliability investments before they become production constraints
• Significant experience in Site Reliability Engineering, performance engineering, platform engineering, distributed systems, or a closely related engineering discipline • Experience supporting production systems with meaningful scale, traffic, latency, or availability requirements • Deep understanding of observability, performance analysis, capacity planning, and reliability engineering • Strong hands-on experience with cloud infrastructure and production distributed systems • Deep knowledge of databases, networking, caching, queueing, compute, storage, and common distributed-system failure modes • Experience defining and operating against SLOs, SLIs, error budgets, and production reliability metrics • Hands-on experience performing load, stress, soak, scalability, and resilience testing • Ability to profile systems, diagnose bottlenecks, tune architecture, and work directly with engineering teams to implement improvements • Experience designing for graceful degradation, dependency failures, recovery, and high-demand scenarios • Strong incident management and root-cause analysis experience • Ability to translate technical performance and reliability risks into clear business implications for senior leadership • Strong judgment regarding optimization versus premature complexity • Nice to have: experience operating high-scale SaaS, identity, DNS, registry, infrastructure, or other highly distributed platforms • Nice to have: experience creating capacity-cost models and forecasting infrastructure requirements • Nice to have: experience building performance and reliability gates into CI/CD pipelines • Nice to have: experience preparing platforms for significant traffic increases associated with enterprise customers or strategic partnerships • Nice to have: experience leading reliability or performance initiatives spanning multiple engineering teams
• Equal Opportunity Employer committed to a diverse and inclusive workplace • Application-process accommodation available through HR
Apply Now