Senior Site Reliability Engineer

🔥 2 hours ago

🏄 California – Remote

info

💵 $158.4k - $294.1k / year

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🦅 H1B Visa Sponsor

info
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Veeam Software

Veeam Software

1001 - 5000 employees

Founded 2006

💼 Consulting

📦 Logistics

☁️ SaaS

💰 $500M Private Equity Round on 2019-01

Consulting • Logistics • SaaS

Veeam Software is a global leader in data resilience and protection, offering self-managed data protection software for hybrid and multi-cloud environments. Their Veeam Data Platform provides comprehensive solutions for data backup, recovery, and security, featuring zero-trust principles and AI-powered tools for data intelligence. Veeam's offerings include secure backup and storage services for platforms such as Microsoft 365, AWS, and Google Cloud, supporting diverse workloads including virtual, physical, and SaaS environments. With a reputation for innovation and customer trust, Veeam serves a broad range of industries, ensuring data resilience against disruptions such as ransomware attacks. Their solutions enable businesses to achieve data freedom, secure storage, and efficient management, reinforcing their position as a top vendor in enterprise backup and recovery software worldwide.

📋 Description

• Build reliability practices for Veeam Data Cloud’s Government and Sovereign Cloud environment • Map platform systems, dependencies, workloads, and risk areas • Create onboarding materials, runbooks, architecture documentation, and operational guides • Design highly available and fault-tolerant infrastructure on Azure and Azure Government • Define SLIs, SLOs, and error budgets • Lead incident response and blameless postmortems • Identify reliability risks and develop remediation plans within compliance constraints • Define observability instrumentation requirements and drive implementation • Establish alerting, telemetry, and monitoring standards • Build automation to reduce toil and support fleet management • Participate in on-call rotations • Work with IaC, CI/CD, deployment automation, and configuration management in air-gapped or restricted environments • Build and maintain testing, canary deployment, and release validation pipelines • Integrate chaos engineering and monitoring tools • Collaborate with product, platform, security, legal, compliance, and operations teams • Own reliability problems end-to-end and drive solutions • Mentor engineers and spread SRE practices across Veeam

🎯 Requirements

• 7+ years in Software Engineering, including 3+ years in SRE, Platform Engineering, or similar, across multi-service platforms • Experience with Government or Sovereign Cloud, such as Azure Government or AWS GovCloud • Experience in regulated compliance environments, including FedRAMP, CMMC, IL2/IL4/IL5, PCI-DSS, SOX, HIPAA, or HITRUST • Strong experience building and running production services on cloud infrastructure; Azure preferred, including Azure Government • Ability to learn large, complex platforms quickly with limited guidance and restricted environment access • Ability to investigate systems independently and produce clear documentation, risk assessments, and improvement plans • Programming skills in TypeScript/JS, Go, Java, C#, or similar • Experience with monitoring and observability tools such as Prometheus, Grafana, OpenTelemetry, or ELK stack • Experience with IaC tools such as Terraform, Terragrunt, or Pulumi • Experience with container orchestration using Kubernetes • Experience with CI/CD and GitOps tooling such as GitHub Actions, Azure DevOps, GitLab CI, ArgoCD, FluxCD, or Dagger • Strong understanding of distributed systems, networking, and cloud-native architecture • Clear written and verbal communication skills • Bonus: experience on B2B SaaS platforms in regulated or government markets • Bonus: background in chaos engineering, resilience testing, or performance/load testing • Bonus: experience building an SRE or reliability function from scratch • Bonus: experience across modern cloud-native and legacy systems • Bonus: familiarity with AI-first development workflows using LLM-powered tools

🏖️ Benefits

• Unlimited paid time off • 12 paid holidays, including 4 global VeeaMe Days for self-care • 24 paid volunteer hours annually through Veeam Cares • Paid parental leave: 8 weeks for all parents, 16 weeks for birthing parents • Medical, dental, and vision coverage starting on the first day • Mental health support, therapy sessions, and digital wellness tools via the Employee Assistance Program • 401(k) retirement plan with company matching contributions • Fertility, adoption, and surrogacy support through Maven • AirVet 24/7 virtual veterinary care at no cost • Legal services, identity protection, and supplemental health insurance options • Tax-advantaged spending accounts for healthcare, dependent care, and commuting • On-demand learning libraries including LinkedIn Learning and O’Reilly • Mentoring, workshops, and learning events including the annual Global Day of Learning • Competitive performance-based bonus included in total target compensation • Comprehensive benefits package including health coverage, retirement plans, and unlimited time off

Apply Now

Similar Jobs

🔥 3 hours ago

Syniti

1001 - 5000

🤝 B2B

🏢 Enterprise

Senior SRE building secure, observable Azure and AWS infrastructure for Syniti’s cloud-native enterprise data platform. Supporting Kubernetes, CI/CD, compliance, and incident response across global SaaS workloads.

🔥 3 hours ago

Karat

201 - 500

👥 HR Tech

🏢 Enterprise

☁️ SaaS

Senior Deployment Engineer helping Karat, a technical interviewing company, implement and optimize enterprise interview frameworks. Advising clients, analyzing hiring performance, and delivering executive training.

🔥 3 hours ago

Cisco

10,000+ employees

🔧 Hardware

🔐 Security

🏢 Enterprise

Site Reliability Engineer architecting Cisco’s developer platform and infrastructure for cloud application delivery. Consolidating legacy tools and improving scalable, resilient engineering workflows.

🔥 3 hours ago

Cisco

10,000+ employees

🔧 Hardware

🔐 Security

🏢 Enterprise

Customer Reliability Engineer resolving escalated Cisco Hypershield incidents across Nexus Smart Switches and on-premises Kubernetes. Improving reliability through diagnostics, runbooks, tooling, and engineering fixes.

🔥 3 hours ago

Fifth Third Bank

10,000+ employees

🛡️ Insurance

💼 Consulting

📦 Logistics

Consulting Site Reliability Engineer designing and supporting complex technology solutions for Fifth Third Bank. Managing vendor software, automation, operational support and engineering guidance.