Senior Platform Engineer

🔥 3 minutes ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Flexential

Flexential

501 - 1000 employees

Founded 2000

🤝 B2B

🏢 Enterprise

📡 Telecommunications

B2B • Enterprise • Telecommunications

Flexential is a US-based provider of data center infrastructure and hybrid IT services. The FlexAnywhere® platform combines colocation, cloud (public/private/hybrid), connectivity/interconnection, data protection, managed and professional services to support enterprise workloads — including high-density GPU and AI/ML deployments. Flexential operates 40+ data centers across 18 US markets (3M+ sq. ft. footprint) and a 100+ Gbps network backbone, and positions itself as a B2B partner for enterprises needing resilient, scalable, and interconnected infrastructure.

📋 Description

• Design, develop and operationally manage automated, resilient, high availability, self-healing, secure platforms with native-AI capabilities for IT needs, serving both internal as well as customer business capabilities. • Develop, and manage the Observability OpenTelemetry Central Backend Stack: Grafana Enterprise, Mimir, Loki, Tempo, and Alertmanager on Kubernetes/RKE2 via Helm and GitLab CI-CD. • Build and manage iaC and CI-CD for automated provisiong and deployment, including Terraform modules for Infra/VM/storage provisioning, Ansible AWX playbooks for OS/App bootstrap, ArgoCD and Helm for Kubernetes configuration. • Develop and manage OpenTelemetry Prometheus scrape profile library including SNMP exporters, REST API exporters, and cloud provider exporters (CloudWatch, Azure Monitor, GCP) for multiple device classes. • Develop AIOps capabilities on platforms for e.g. Observability use-cases: anomaly detection integrations, event correlation rules in Alertmanager, and synthetic monitoring patterns to reduce alert noise. • Configure and maintain Zabbix auto-discovery: network range scanning, device classification, and Prometheus service discovery integration. • Build and harden Edge Stack deployments (Prometheus + OTel collector) per data center site using GitOps templates. • Integrate Alertmanager with ServiceNow: webhook routing, ticket enrichment, auto-close logic, and escalation policy configuration. • Maintain platform security: Conjur/CyberArk secret injection at runtime, mTLS between stack components, RBAC in Grafana Enterprise. • Author and maintain Grafana dashboards in JSON/GitLab — facility overview, network health, RED metrics, application telemetry. • Mentor mid-level engineers, lead code reviews, and establish engineering standards for the team. • Represent platform engineering in cross-functional architecture reviews and executive-level program updates. • Perform other duties as required and assigned.

🎯 Requirements

• 5+ years in a production environment. • Kubernetes (RKE2/k3s). • Helm chart deployment. • systemd services. • Docker/containerd. • 4+ years: Grafana, Mimir, Loki, Tempo configuration, tuning, dash-boarding and production operations. • Prometheus required. • 5+ years Senior-Level Python / Scripting Frameworks. • Automation scripts. • Exporter development. • GitLab pipeline scripting. • REST API integrations. • 5+ years GitOps / CI/CD. • GitLab CI/CD pipeline authoring. • Terraform and Ansible as primary IaC tools. • ArgoCD or Flux preferred. • 2+ years AIOps / Observability Engineering. • Alertmanager rule authoring. • Anomaly detection integration. • Event correlation. • Noise reduction techniques. • 5+ years Working Infrastructure (Linux/VM) Management Knowledge. • Linux administration. • VMware vCenter/VCF experience. • Netapp storage management. • Network fundamentals (SNMP, TCP/IP). • 2+ years Secrets Management. • CyberArk/Conjur, HashiCorp Vault, or equivalent. • Runtime secret injection patterns. • Minimal travel may be required.

🏖️ Benefits

• Medical, Telehealth, Dental and Vision • 401(k) • Health Savings Accounts (HSA) and Flexible Spending Accounts (FSA) • Life and AD&D • Short Term and Long-Term disability • Flex Paid Time Off (PTO) • Leave of Absence • Employee Assistance Program • Wellness Program • Rewards and Recognition Program

Apply Now

Similar Jobs

🔥 3 hours ago

Consensus Cloud Solutions

501 - 1000

🏥 Healthcare

☁️ SaaS

🤖 Artificial Intelligence

Technical leader driving implementation of eCommerce and customer experience platform at Consensus Cloud Solutions. Shifting engineering strategy to a flexible, scalable, distributed web ecosystem.

🇺🇸 United States – Remote

💵 $140k - $197.5k / year

💰 $225M Post-IPO Debt - Consensus Cloud Solutions on 2025-07

⏰ Full Time

🟠 Senior

🏗️ Platform Engineer

Distributed Systems

Java

React

Spring

SQL

🔥 3 hours ago

iFIT

1001 - 5000

🧘 Wellness

☁️ SaaS

🔧 Hardware

Senior Software Engineer developing data platforms at iFIT, integrating AI into fitness experiences. Responsible for building backend infrastructure and enhancing real-time data processing.

🇺🇸 United States – Remote

💵 $130k - $160k / year

💰 $200M Private Equity Round - iFit on 2019-12

⏰ Full Time

🟠 Senior

🏗️ Platform Engineer

AWS

GraphQL

JavaScript

Node.js

TypeScript

🔥 3 hours ago

iFIT

1001 - 5000

🧘 Wellness

☁️ SaaS

🔧 Hardware

Tech Lead/Sr. Manager of Platform Engineering at iFIT responsible for cloud ops, infrastructure, and AI innovation. Leading technical roadmap and ensuring reliability in platform engineering.

🇺🇸 United States – Remote

💵 $190k - $225k / year

💰 $200M Private Equity Round - iFit on 2019-12

⏰ Full Time

🟡 Mid-level

🟠 Senior

🏗️ Platform Engineer

AWS

Azure

Cloud

Google Cloud Platform

Terraform

🔥 8 hours ago

Aalyria

51 - 200

📡 Telecommunications

🏢 Enterprise

☁️ SaaS

Software Engineer working on AI products and platform services for aerospace technology. Responsible for building, operating, and improving AI systems and integrations within engineering workflows.

Cloud

Kubernetes

Python

Terraform

TypeScript

Go

🔥 9 hours ago

INTERNATIONAL RESCUE COMMITTEE

10,000+ employees

🏥 Healthcare

📦 Logistics

🤲 Charity

AI Platform Engineer managing the technical operations of IRC's enterprise AI environments and integrations, ensuring reliability and performance while collaborating with engineering teams.

Azure

Cloud

Distributed Systems

JavaScript

Python

TypeScript