Vice President, Site Reliability Engineering – Data Centers

🕒 6 days ago

🇺🇸 United States – Remote

⏰ Full Time

🔴 Lead

⛑ DevOps & Site Reliability Engineer (SRE)

🦅 H1B Visa Sponsor

info
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Galaxy

Galaxy

201 - 500 employees

Founded 2018

₿ Crypto

💸 Finance

Crypto • Finance

Galaxy is a digital asset and blockchain leader helping institutions, startups, and qualified individuals shape a changing economy through innovative crypto solutions. Galaxy provides a wide array of services, including asset management, trading, lending, custodial technology, and blockchain infrastructure solutions. With a focus on both traditional finance integration and digital asset expertise, Galaxy is committed to advancing the adoption and functionality of cryptocurrencies and blockchain technologies across the globe.

📋 Description

• Oversee an SRE team focused on designing, deploying, and maintaining automation toolsets and related systems • Establish and enforce Infrastructure as Code standards for consistent, repeatable, and secure deployments • Lead automated configuration and state management using Ansible playbooks and Packer image pipelines across Windows, Linux, and ESXi platforms • Manage monitoring and health of automation platforms and implement SLIs/SLOs • Drive automated lifecycle management of physical and virtual assets, including template creation, deployment, patching, scaling, and decommissioning • Lead development of custom scripts and internal providers using Python, Go, PowerShell, and Bash • Collaborate with the broader Datacenter team and facilitate team-wide workflows • Analyze system behavior and resource utilization in virtual environments to optimize automated deployment performance • Provide technical guidance and career mentorship to SREs and foster an automate-first culture

🎯 Requirements

• 6–10 years’ experience in Infrastructure, SRE, or DevOps focused on infrastructure automation at scale • Deep proficiency with Terraform, including providers, modules, and state management • Deep proficiency with Ansible, including roles, playbooks, and Tower/AWX • Hands-on experience creating standardized, hardened Windows and Linux images using Packer, Ansible, or SCCM • Strong experience managing and automating VMware vSphere/vCenter, Azure, and AWS • High-level scripting skills in Python, Go, PowerShell, and Bash • Experience with Splunk, ELK, Prometheus, or Grafana for infrastructure health and automation telemetry • Understanding of network topology and design, including Juniper Networks or Palo Alto platforms • Strong mastery of Git, including branching strategies and pull-request workflows • Experience with CI/CD platforms such as Jenkins, GitLab CI, or GitHub Actions • Comfort managing, troubleshooting, and tuning both Windows Server and Linux • Previous team leadership or management experience • Experience with IAM platforms such as Entra ID, Active Directory, or Okta • Experience with block- and object-based storage solutions on-premises or in the cloud • Storage backup/disaster recovery administration with Commvault or Veeam

🏖️ Benefits

• Equal employment opportunities • Reasonable accommodation for qualified applicants with disabilities

Apply Now

Similar Jobs

🕒 6 days ago

Intus Care

11 - 50

💼 Consulting

📣 Marketing

📦 Logistics

Director of SRE leading reliability, QA, observability, and incident management for Intus Care’s cloud-native healthcare EMR platform. Building scalable SRE capabilities and operational standards for systems supporting value-based care.

🇺🇸 United States – Remote

💵 $175k - $200k / year

💰 $13.1M Venture Round on 2023-01

⏰ Full Time

🔴 Lead

⛑ DevOps & Site Reliability Engineer (SRE)

🦅 H1B Visa Sponsor

info

🕒 6 days ago

Veeam Software

1001 - 5000

💼 Consulting

📦 Logistics

☁️ SaaS

Site Reliability Engineer building reliability practices for Veeam’s Government and Sovereign Cloud SaaS platform. Designing Azure infrastructure, observability, automation, and incident-response systems in regulated environments.

🇺🇸 United States – Remote

💵 $138.9k - $231.4k / year

💰 $500M Private Equity Round on 2019-01

⏰ Full Time

🟠 Senior

🔴 Lead

⛑ DevOps & Site Reliability Engineer (SRE)

🦅 H1B Visa Sponsor

info

🕒 August 1

Filevine

201 - 500

☁️ SaaS

⚖️ Legal

🤖 Artificial Intelligence

Staff Site Reliability Engineer at Filevine shaping engineering culture and driving reliability practices. Leading technical standards and mentorship within a remote engineering team focused on legal AI technology.

🇺🇸 United States – Remote

💵 $235k - $275k / year

💰 $108M Series D on 2022-04

⏰ Full Time

🔴 Lead

⛑ DevOps & Site Reliability Engineer (SRE)

🦅 H1B Visa Sponsor

info

🕒 July 31

Aya Healthcare

5001 - 10000

🏥 Healthcare

💼 Consulting

📦 Logistics

Manager of Site Reliability Engineering leading a team for Aya Healthcare's workforce platform. Ensuring product reliability and outstanding user experience through innovative solutions.

🕒 July 30

TalentWerx

11 - 50

🎯 Recruiter

👥 HR Tech

🤝 B2B

DevOps Engineer IV designing and optimizing deployment solutions for Aether Aerospace. Collaborating with developers to enhance software development processes and ensure system security.

🇺🇸 United States – Remote

💵 $123.6k - $159k / year

⏰ Full Time

🟠 Senior

🔴 Lead

⛑ DevOps & Site Reliability Engineer (SRE)