
11 - 50 Mitarbeiter
🔧 Hardware
🏢 Unternehmen
🤖 Künstliche Intelligenz
💰 €10.000.000 Seed Round im 2022-04
Hardware • Enterprise • Artificial Intelligence
Hydra Host ist ein Anbieter von Hochleistungs-Computing-Lösungen und bietet dedizierten Bare-Metal-GPU-Serverzugang, der für KI- und HPC-Workloads optimiert ist. Ihre Plattform ermöglicht es Benutzern, weltweit auf erstklassige GPUs zuzugreifen und diese zu mieten, wobei unübertroffene Leistung, Sicherheit und Anpassung geboten werden. Die Infrastruktur von Hydra Host umfasst einen Marktplatz namens Brokkr, der eine breite Palette von GPU-Konfigurationen und Lösungen für geschäftskritische Anwendungen wie KI, Big Data und maschinelles Lernen bietet. Durch ihre robusten, sicheren und skalierbaren Lösungen stellt Hydra Host sicher, dass Kunden die volle Kontrolle über ihre Serverumgebungen genießen können, mit Optionen zur Skalierbarkeit und Zukunftssicherheit. Die Angebote des Unternehmens werden von führenden Unternehmen geschätzt, die effiziente und innovative Computing-Lösungen suchen.
🕒 vor 9 Monaten
🐊 Florida – Remote
💵 $140.000 - $200.000 / Jahr
⏰ Vollzeit
🟡 Mittelstufe
🟠 Senior
⛑ DevOps- und Site Reliability Engineer (SRE)
👻 Geisterscore 58%
🗣️🇺🇸🇬🇧 Englisch erforderlich
Verbessern Sie Ihre Chancen auf ein Vorstellungsgespräch, indem Sie Ihre Lebenslauf-Bewertung vor der Bewerbung überprüfen.

11 - 50 Mitarbeiter
🔧 Hardware
🏢 Unternehmen
🤖 Künstliche Intelligenz
💰 €10.000.000 Seed Round im 2022-04
Hardware • Enterprise • Artificial Intelligence
Hydra Host ist ein Anbieter von Hochleistungs-Computing-Lösungen und bietet dedizierten Bare-Metal-GPU-Serverzugang, der für KI- und HPC-Workloads optimiert ist. Ihre Plattform ermöglicht es Benutzern, weltweit auf erstklassige GPUs zuzugreifen und diese zu mieten, wobei unübertroffene Leistung, Sicherheit und Anpassung geboten werden. Die Infrastruktur von Hydra Host umfasst einen Marktplatz namens Brokkr, der eine breite Palette von GPU-Konfigurationen und Lösungen für geschäftskritische Anwendungen wie KI, Big Data und maschinelles Lernen bietet. Durch ihre robusten, sicheren und skalierbaren Lösungen stellt Hydra Host sicher, dass Kunden die volle Kontrolle über ihre Serverumgebungen genießen können, mit Optionen zur Skalierbarkeit und Zukunftssicherheit. Die Angebote des Unternehmens werden von führenden Unternehmen geschätzt, die effiziente und innovative Computing-Lösungen suchen.
• Design, deploy, and maintain QA systems used by our development teams to test integration and live system responses across full-stack deployments in local, live, and ephemeral environments • Evaluate and integrate monitoring and QA tools to find the right tools for the job • Create a unified monitoring platform and processes that datacenter and device teams will integrate to monitor their components (live servers, lifecycle, networks, power, etc.) • Maintain monitoring processes and dashboards to provide complete visibility into the health, performance, and reliability of our CI systems, software deployments, and testing platforms • Create and maintain a systems test suite, in collaboration with our product managers, to validate marketplace changes against all business functions in live and ephemeral QA environments • Integrate all fore-mentioned systems to create holistic platform health statistics reporting • Design disaster-recovery processes in collaboration with devops • Ensure we are meeting uptime SLAs across all platform deployments • Work with datacenter and device teams to define service-level indicators (SLIs), service-level objectives (SLOs), and SLAs • Establish observability standards across the stack: logs, metrics, traces, and alerts, and actionable on-call playbooks • Automate everything from monitoring setups to incident responses to eliminate manual toil and increase reliability • Drive incident response, root cause analysis, and post‑mortems • Guide incident turn-around into tooling and process improvements • Establish the monitoring infrastructure and dashboards that enable everyone — from engineers to execs — to know what’s going on • Act as the reliability partner to engineering teams: review systems for reliability concerns, help design QA requirements and testing, and help teams meet reliability targets.
• 5–8+ years of experience in Reliability Engineering, DevOps, or infrastructure roles focused on large-scale, high-uptime production environments • Deep familiarity with monitoring and observability tooling: you've implemented and managed systems, esp. Prometheus, Grafana, and Zabbix • Strong experience with service orchestration in mutli-region environment (Nomad, Kubernetes, cloud VMs, distributed databases) • Track record of managing production system uptime and SLAs and building tools to support it • Experience writing and reviewing post-mortems and using those findings to drive improvements in tools and process • Proficient with scripting and programming languages (Python, Go, BASH, etc.) for automating operational tasks • Strong proficiency with infrastructure as code and devops workflows • Experience with distributed tracing, log aggregation, and alert tuning • Passion for building systems that fail gracefully, alert correctly, and empower others to operate confidently • Excellent communication skills: you can write clear documentation, drive incident reviews, and communicate reliability risks to technical and non-technical stakeholders.
• Competitive compensation: base salary + performance bonus + equity • Exposure to high-performance computing and state-of-the-art GPU environments • A core role in ensuring our systems are reliable, observable, and meet customer SLAs • Remote work environment with a strong culture of ownership and autonomy • No red tape: find the right solution, work with the team, get feedback, and get the job done.
Jetzt Bewerben🕒 vor 10 Monaten
CloudOps & DevOps Engineer enhancing secure data flows between Kafka, PostgreSQL, and S3. Collaborating with teams to ensure reliable infrastructure for client projects.
🇺🇸 Vereinigte Staaten – Remote
⏰ Vollzeit
🟡 Mittelstufe
🟠 Senior
⛑ DevOps- und Site Reliability Engineer (SRE)
🦅 H1B-Visum-Sponsor
🗣️🇺🇸🇬🇧 Englisch erforderlich
🕒 vor 10 Monaten
Talent-pool for DevOps-specialist roles at Mission Box Solutions. Connecting veteran-owned recruiting agency candidates with hiring companies across DevOps specializations.
🇺🇸 Vereinigte Staaten – Remote
⏰ Vollzeit
🟡 Mittelstufe
🟠 Senior
⛑ DevOps- und Site Reliability Engineer (SRE)
🗣️🇺🇸🇬🇧 Englisch erforderlich
🕒 vor 11 Monaten
Remote Site Reliability Engineer at ContainIQ maintaining cloud-native observability platform; job description coming soon; contact careers email.
🇺🇸 Vereinigte Staaten – Remote
💰 €2.500.000 Seed Round im 2021-10
⏰ Vollzeit
🟡 Mittelstufe
🟠 Senior
⛑ DevOps- und Site Reliability Engineer (SRE)
🗣️🇺🇸🇬🇧 Englisch erforderlich
🕒 vor 11 Monaten
Senior DevOps Engineer building secure, scalable AWS infrastructure for an AI brand-safety contextual intelligence platform. Lead CI/CD, serverless, observability, and incident response.
🗣️🇺🇸🇬🇧 Englisch erforderlich
🕒 vor 12 Monaten
Expression-of-interest for Site Reliability Engineer at GitLab. Remote pipeline role for candidates to join AI-powered DevSecOps teams.
🇺🇸 Vereinigte Staaten – Remote
⏰ Vollzeit
🟡 Mittelstufe
🟠 Senior
⛑ DevOps- und Site Reliability Engineer (SRE)
🗣️🇺🇸🇬🇧 Englisch erforderlich