Staff Site Reliability Engineer

🔥 3 minutes ago

🇪🇸 Spain – Remote

⏰ Full Time

🔴 Lead

⛑ DevOps & Site Reliability Engineer (SRE)

👻 Ghost score 10%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Levi Strauss & Co.

Levi Strauss & Co.

10,000+ employees

Founded 1873

🏭 Manufacturing

💼 Consulting

👗 Fashion

Manufacturing • Consulting • Fashion

Levi Strauss & Co. is a leading company in the apparel and fashion industry, known for its iconic denim jeans. Founded in 1853, Levi's has become a global brand with a rich heritage in crafting durable and stylish clothing. It is recognized for innovation in denim and its commitment to sustainability and ethical practices in its manufacturing processes. Levi Strauss & Co. continues to evolve, offering a diverse range of fashion products to meet the needs of modern consumers worldwide.

📋 Description

• Own and elevate the reliability, scalability, and operability of enterprise data and AI platforms • Define, instrument, and enforce SLOs, SLIs, and error budgets across platform services • Reduce MTTD and MTTR through improved observability, automated alerting, and runbook-driven incident response • Lead blameless post-mortems and implement durable reliability improvements • Identify, measure, and eliminate operational toil; track toil percentage and enforce capacity guardrails • Build and maintain self-service infrastructure capabilities for product and data engineering teams • Automate deployment pipelines, configuration management, and operational workflows using Infrastructure-as-Code, Terraform, Helm, and GitOps • Architect and optimize workloads across GCP services including GKE, Cloud Run, BigQuery, Pub/Sub, GCS, Composer, Dataflow, and Vertex AI • Lead multi-cloud architecture decisions across GCP and Azure • Design self-healing infrastructure, auto-scaling strategies, and capacity planning models • Champion data security and governance practices including encryption, least-privilege IAM, secrets management, and audit logging • Apply SRE principles to agentic AI workloads and define reliability expectations for LLM-based and multi-agent systems • Partner with AI Platform teams to productionize agentic pipelines with monitoring, drift detection, and rollback capabilities • Drive adoption of AI-assisted operations tooling for observability, anomaly detection, and predictive incident management • Guide and mentor junior and mid-level SREs through code reviews and reliability reviews • Collaborate with Data Engineering, Software Engineering, Security, and Product teams • Communicate platform health, risk posture, and reliability roadmaps to technical and executive audiences

🎯 Requirements

• Master's degree in Computer Science, Engineering, or related field (or equivalent practical experience) • 10+ years of experience in Site Reliability Engineering, DevOps, or Platform Engineering, with a strong track record in large-scale production environments • Deep, hands-on expertise in GCP, including GKE, Cloud Run, BigQuery, Pub/Sub, GCS, Composer, Dataflow, and Vertex AI • Proficiency with Infrastructure-as-Code tools (Terraform, Helm) and GitOps workflows (ArgoCD, Flux) • Strong command of observability tooling, including distributed tracing, structured logging, metrics pipelines, and alerting platforms such as Cloud Monitoring, Datadog, and Prometheus/Grafana • Proven experience defining and operating against SLOs, SLIs, and error budgets in production environments • Solid understanding of data security principles: IAM, encryption, secrets management, network policies, and compliance frameworks • Experience with multi-cloud environments (GCP + Azure), including cross-cloud networking, identity federation, and cost governance • Demonstrated ability to lead without authority, influence engineers across teams, and drive reliability improvements organization-wide • Excellent written and verbal communication skills; ability to translate complex reliability concepts for non-technical stakeholders • Fluency in at least one systems or scripting language: Python, Go, or Bash • Experience with container orchestration (Kubernetes/GKE), service mesh, and traffic management patterns • Familiarity with data engineering patterns: batch and streaming pipelines, data warehouses, and operational challenges of large-scale data platforms • Understanding of agentic AI architectures and reliability challenges of LLM-based, event-driven, and multi-agent systems • Working knowledge of data governance frameworks, data lineage tooling, and platform-level data quality enforcement • Experience operating data platforms in retail or e-commerce environments • Familiarity with SRE principles in practice, including error budget policies, CRE engagements, and production readiness reviews • Exposure to FinOps practices, including cloud cost attribution, commitment optimization, and unit economics for data workloads • Experience with Azure-native services (AKS, Azure Data Factory, Event Hubs) and cross-cloud identity management • Prior experience in a Staff or Principal-level SRE role with organization-wide scope

Apply Now

Similar Jobs

🕒 August 21

MEDvidi

201 - 500

🏥 Healthcare

⚕️ Healthcare Insurance

Staff DevOps Engineer owning AWS, Kubernetes, security, and developer experience. Scaling MEDvidi’s AI-powered psychiatric care platform for safe, effective psychiatric treatment.

🗣️🇷🇺 Russian Required

AWS

DNS

Docker

EC2

Grafana

Kubernetes

Linux

Node.js

Postgres

Prometheus

Terraform

TypeScript

Vault

🕒 May 12

Stellar Cyber

51 - 200

🔒 Cybersecurity

🤖 Artificial Intelligence

🏢 Enterprise

Seeking a Staff Site Reliability Engineer to enhance scalability and efficiency for production systems at Stellar Cyber. Join a leader in cybersecurity focused on innovative solutions against cyber threats.

AWS

Azure

Cloud

Distributed Systems

ElasticSearch

Google Cloud Platform

Grafana

Kafka

Kubernetes

Linux

MongoDB

Prometheus

Python

Redis

Spark

Terraform