Senior Platform Engineer

🕒 vor 4 Tagen

🗣️🇺🇸🇬🇧 Englisch erforderlich

Jetzt Bewerben
Ähnliche Remote-Jobs finden

📊 Überprüfen Sie Ihre Lebenslauf-Bewertung für diese Stelle

Verbessern Sie Ihre Chancen auf ein Vorstellungsgespräch, indem Sie Ihre Lebenslauf-Bewertung vor der Bewerbung überprüfen.

Logo of Flexential

Flexential

501 - 1000 Mitarbeiter

Gegründet 2000

🤝 B2B

🏢 Unternehmen

📡 Telekommunikation

B2B • Enterprise • Telecommunications

Flexential ist ein in den USA ansässiger Anbieter von Rechenzentrumsinfrastruktur und hybriden IT-Diensten. Die FlexAnywhere®-Plattform vereint Colocation, Cloud (öffentlich/privat/hybrid), Konnektivität/Interkonnektion, Datenschutz, Managed- und Professional Services, um Unternehmens-Workloads zu unterstützen – einschließlich hochdichter GPU- und KI/ML-Implementierungen. Flexential betreibt mehr als 40 Rechenzentren in 18 US-Märkten (mehr als 3 Millionen Quadratfuß Fläche) und ein Netzwerk-Backbone mit über 100 Gbps und positioniert sich als B2B-Partner für Unternehmen, die eine widerstandsfähige, skalierbare und vernetzte Infrastruktur benötigen.

Beschreibung

• Design, develop and operationally manage automated, resilient, high availability, self-healing, secure platforms with native-AI capabilities for IT needs, serving both internal as well as customer business capabilities. • Develop, and manage the Observability OpenTelemetry Central Backend Stack: Grafana Enterprise, Mimir, Loki, Tempo, and Alertmanager on Kubernetes/RKE2 via Helm and GitLab CI-CD. • Build and manage iaC and CI-CD for automated provisiong and deployment, including Terraform modules for Infra/VM/storage provisioning, Ansible AWX playbooks for OS/App bootstrap, ArgoCD and Helm for Kubernetes configuration. • Develop and manage OpenTelemetry Prometheus scrape profile library including SNMP exporters, REST API exporters, and cloud provider exporters (CloudWatch, Azure Monitor, GCP) for multiple device classes. • Develop AIOps capabilities on platforms for e.g. Observability use-cases: anomaly detection integrations, event correlation rules in Alertmanager, and synthetic monitoring patterns to reduce alert noise. • Configure and maintain Zabbix auto-discovery: network range scanning, device classification, and Prometheus service discovery integration. • Build and harden Edge Stack deployments (Prometheus + OTel collector) per data center site using GitOps templates. • Integrate Alertmanager with ServiceNow: webhook routing, ticket enrichment, auto-close logic, and escalation policy configuration. • Maintain platform security: Conjur/CyberArk secret injection at runtime, mTLS between stack components, RBAC in Grafana Enterprise. • Author and maintain Grafana dashboards in JSON/GitLab — facility overview, network health, RED metrics, application telemetry. • Mentor mid-level engineers, lead code reviews, and establish engineering standards for the team. • Represent platform engineering in cross-functional architecture reviews and executive-level program updates. • Perform other duties as required and assigned.

🎯 Anforderungen

• 5+ years in a production environment. • Kubernetes (RKE2/k3s). • Helm chart deployment. • systemd services. • Docker/containerd. • 4+ years: Grafana, Mimir, Loki, Tempo configuration, tuning, dash-boarding and production operations. • Prometheus required. • 5+ years Senior-Level Python / Scripting Frameworks. • Automation scripts. • Exporter development. • GitLab pipeline scripting. • REST API integrations. • 5+ years GitOps / CI/CD. • GitLab CI/CD pipeline authoring. • Terraform and Ansible as primary IaC tools. • ArgoCD or Flux preferred. • 2+ years AIOps / Observability Engineering. • Alertmanager rule authoring. • Anomaly detection integration. • Event correlation. • Noise reduction techniques. • 5+ years Working Infrastructure (Linux/VM) Management Knowledge. • Linux administration. • VMware vCenter/VCF experience. • Netapp storage management. • Network fundamentals (SNMP, TCP/IP). • 2+ years Secrets Management. • CyberArk/Conjur, HashiCorp Vault, or equivalent. • Runtime secret injection patterns. • Minimal travel may be required.

🏖️ Vorteile

• Medical, Telehealth, Dental and Vision • 401(k) • Health Savings Accounts (HSA) and Flexible Spending Accounts (FSA) • Life and AD&D • Short Term and Long-Term disability • Flex Paid Time Off (PTO) • Leave of Absence • Employee Assistance Program • Wellness Program • Rewards and Recognition Program

Jetzt Bewerben

Ähnliche Jobs

🕒 vor 4 Tagen

Consensus Cloud Solutions

501 - 1000

🏥 Gesundheitswesen

☁️ SaaS

🤖 Künstliche Intelligenz

Technical leader driving implementation of eCommerce and customer experience platform at Consensus Cloud Solutions. Shifting engineering strategy to a flexible, scalable, distributed web ecosystem.

🇺🇸 Vereinigte Staaten – Remote

💵 $140.000 - $197.500 / Jahr

💰 €225.000.000 Post-IPO Debt - Consensus Cloud Solutions im 2025-07

⏰ Vollzeit

🟠 Senior

🏗️ Plattformingenieur

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 4 Tagen

iFIT

1001 - 5000

🧘 Wellness

☁️ SaaS

🔧 Hardware

Senior Software Engineer developing data platforms at iFIT, integrating AI into fitness experiences. Responsible for building backend infrastructure and enhancing real-time data processing.

🇺🇸 Vereinigte Staaten – Remote

💵 $130.000 - $160.000 / Jahr

💰 €200.000.000 Private Equity Round - iFit im 2019-12

⏰ Vollzeit

🟠 Senior

🏗️ Plattformingenieur

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 4 Tagen

iFIT

1001 - 5000

🧘 Wellness

☁️ SaaS

🔧 Hardware

Tech Lead/Sr. Manager of Platform Engineering at iFIT responsible for cloud ops, infrastructure, and AI innovation. Leading technical roadmap and ensuring reliability in platform engineering.

🇺🇸 Vereinigte Staaten – Remote

💵 $190.000 - $225.000 / Jahr

💰 €200.000.000 Private Equity Round - iFit im 2019-12

⏰ Vollzeit

🟡 Mittelstufe

🟠 Senior

🏗️ Plattformingenieur

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 5 Tagen

Aalyria

51 - 200

📡 Telekommunikation

🏢 Unternehmen

☁️ SaaS

Software Engineer working on AI products and platform services for aerospace technology. Responsible for building, operating, and improving AI systems and integrations within engineering workflows.

🇺🇸 Vereinigte Staaten – Remote

💵 $185.000 - $215.000 / Jahr

⏰ Vollzeit

🟡 Mittelstufe

🟠 Senior

🏗️ Plattformingenieur

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 5 Tagen

The Home Depot

10.000+ Mitarbeiter

🏗️ Bauwesen

📦 Logistik

🛒 Einzelhandel

Engineering Manager leading a high-performing team building foundational customer platforms for The Home Depot. Overseeing engineering delivery and collaborating with product and engineering teams.

🇺🇸 Vereinigte Staaten – Remote

💵 $140.000 - $240.000 / Jahr

💰 Debt Financing im 2007-07

⏰ Vollzeit

🟡 Mittelstufe

🟠 Senior

🏗️ Plattformingenieur

🗣️🇺🇸🇬🇧 Englisch erforderlich