SRE Architect, AI-Powered Reliability

🕒 May 13

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of WEX

WEX

5001 - 10000 employees

Founded 1983

🏥 Healthcare

📦 Logistics

✈️ Travel

💰 $310M Post-IPO Debt on 2020-06

Healthcare • Logistics • Travel

WEX is a global commerce platform specializing in various business solutions to address operational challenges. They provide services in managing and mobilizing fleets with their fuel card systems, offering comprehensive fleet management and analytics. Additionally, they focus on business payments solutions that streamline processes across industries, enhancing efficiency and security. WEX is also involved in employee benefits administration, helping organizations effectively manage health and reimbursement accounts. Their diverse range of services caters to numerous sectors, emphasizing innovation, sustainability, and effective solutions for business growth.

📋 Description

• Define, publish, and enforce enterprise-wide SRE best practices and operational standards • Define and lead WEX’s AI-Powered Reliability Engineering strategy • Architect and oversee the implementation of mission-critical systems • Establish and govern SLO, SLI, and error budget frameworks across LOBs • Own the production readiness review process • Serve as the primary technical advisor to engineering leadership

🎯 Requirements

• 12+ years in SRE, platform engineering, or distributed systems • Deep practical expertise across observability, incident management, resilience engineering, and capacity planning • Experience with high-availability, low-latency systems • Demonstrated experience using AI tools to solve real reliability problems • Proven ability to define and enforce technical standards across multiple engineering teams • Experience designing self-healing and auto-recovery mechanisms in production distributed systems • Strong background in cloud cost optimization and tooling for managing cloud spend at scale • Excellent written and verbal communication skills

🏖️ Benefits

• health, dental and vision insurances • retirement savings plan • paid time off • health savings account • flexible spending accounts • life insurance • disability insurance • tuition reimbursement • comprehensive and market competitive benefits

Apply Now

Similar Jobs

🕒 May 13

FICO

1001 - 5000

💼 Consulting

🛡️ Insurance

🏥 Healthcare

Senior DevOps Engineer with extensive Kubernetes and AWS experience at FICO. Responsible for CI/CD pipeline architecture and cloud security compliance.

🇺🇸 United States – Remote

💵 $115.5k - $181.5k / year

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🕒 May 11

PlayOn! Sports

201 - 500

📣 Marketing

📦 Logistics

💼 Consulting

Senior Site Reliability Engineer focused on building tools and automation for system reliability at PlayOn. Collaborating with DevOps and engineering teams to enhance performance and scalability.

🇺🇸 United States – Remote

💰 $26M Series D on 2013-07

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🕒 May 9

Visionary Integration Professionals (VIP)

501 - 1000

💼 Consulting

🎖️ Defense

🏥 Healthcare

Forward Deployment Engineer working on AI-enabled solutions for clients at Visionary Integration Professionals. Collaborating with customers to design, prototype, and support implementations across various sectors.

🇺🇸 United States – Remote

💵 $130k - $165k / year

💰 Debt Financing on 2018-11

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🕒 May 8

Zippy

51 - 200

🛡️ Insurance

💸 Finance

💳 Fintech

Director of DevOps overseeing IT and security for a remote-first lending company dedicated to manufactured home loans. Leading high-performing teams and strategic initiatives across technology operations.

🇺🇸 United States – Remote

⏰ Full Time

🔴 Lead

⛑ DevOps & Site Reliability Engineer (SRE)

🕒 May 8

Zingtree

11 - 50

🏥 Healthcare

🛡️ Insurance

📦 Logistics

Senior DevOps / Platform Reliability Engineer managing CI/CD, infrastructure for AI-driven platform. Collaborating across teams to automate processes and ensure reliability.