Senior Platform Engineer

🔥 0 minutes ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Flexential

Flexential

501 - 1000 employees

Founded 2000

🤝 B2B

🏢 Enterprise

📡 Telecommunications

B2B • Enterprise • Telecommunications

Flexential is a US-based provider of data center infrastructure and hybrid IT services. The FlexAnywhere® platform combines colocation, cloud (public/private/hybrid), connectivity/interconnection, data protection, managed and professional services to support enterprise workloads — including high-density GPU and AI/ML deployments. Flexential operates 40+ data centers across 18 US markets (3M+ sq. ft. footprint) and a 100+ Gbps network backbone, and positions itself as a B2B partner for enterprises needing resilient, scalable, and interconnected infrastructure.

📋 Description

• Design, develop and operationally manage automated, resilient, high availability, self-healing, secure platforms with native-AI capabilities for IT needs, serving both internal as well as customer business capabilities. • Develop, and manage the Observability OpenTelemetry Central Backend Stack: Grafana Enterprise, Mimir, Loki, Tempo, and Alertmanager on Kubernetes/RKE2 via Helm and GitLab CI-CD. • Build and manage iaC and CI-CD for automated provisiong and deployment, including Terraform modules for Infra/VM/storage provisioning, Ansible AWX playbooks for OS/App bootstrap, ArgoCD and Helm for Kubernetes configuration. • Develop and manage OpenTelemetry Prometheus scrape profile library including SNMP exporters, REST API exporters, and cloud provider exporters (CloudWatch, Azure Monitor, GCP) for multiple device classes. • Develop AIOps capabilities on platforms for e.g Observability use-cases: anomaly detection integrations, event correlation rules in Alertmanager, and synthetic monitoring patterns to reduce alert noise. • Configure and maintain Zabbix auto-discovery: network range scanning, device classification, and Prometheus service discovery integration. • Build and harden Edge Stack deployments (Prometheus + OTel collector) per data center site using GitOps templates. • Integrate Alertmanager with ServiceNow: webhook routing, ticket enrichment, auto-close logic, and escalation policy configuration. • Maintain platform security: Conjur/CyberArk secret injection at runtime, mTLS between stack components, RBAC in Grafana Enterprise. • Author and maintain Grafana dashboards in JSON/GitLab — facility overview, network health, RED metrics, application telemetry. • Mentor mid-level engineers, lead code reviews, and establish engineering standards for the team. • Represent platform engineering in cross-functional architecture reviews and executive-level program updates. • Perform other duties as required and assigned.

🎯 Requirements

• 5+ years in a production environment • Kubernetes (RKE2/k3s) • Helm chart deployment • System services • Docker/container • 4+ years: Grafana, Mimir, Loki, Tempo configuration, tuning, dash-boarding and production operations • Prometheus required • 5+ years senior-level Python / Scripting frameworks • Automation scripts • Exporter development • GitLab pipeline scripting • REST API integrations • 5+ years GitOps / CI/CD • GitLab CI/CD pipeline authoring • Terraform and Ansible as primary IaC tools • ArgoCD or Flux preferred • 2+ years AIOps / Observability Engineering • Alertmanager rule authoring • Anomaly detection integration • Event correlation • Noise reduction techniques • 5+ years Working Infrastructure (Linux/VM) Management Knowledge • Linux administration • VMware vCenter/VCF experience • Netapp storage management • Network fundamentals (SNMP, TCP/IP) • 2+ years Secrets Management • CyberArk/Conjur, HashiCorp Vault, or equivalent • Runtime secret injection patterns • Minimal travel may be required

🏖️ Benefits

• Medical, Telehealth, Dental and Vision • 401(k) • Health Savings Accounts (HSA) and Flexible Spending Accounts (FSA) • Life and AD&D • Short Term and Long-Term disability • Flex Paid Time Off (PTO) • Leave of Absence • Employee Assistance Program • Wellness Program • Rewards and Recognition Program

Apply Now

Similar Jobs

🔥 55 minutes ago

Quantiphi

1001 - 5000

💼 Consulting

🏥 Healthcare

📦 Logistics

Senior Platform Engineer at Quantiphi designing and optimizing infrastructure for GenAI and LLM workloads. Collaborating with data science and application teams to deliver AI solutions.

Ansible

AWS

Azure

Cloud

Google Cloud Platform

Kubernetes

Linux

OpenShift

Terraform

🔥 1 hour ago

Availity

1001 - 5000

🏥 Healthcare

💼 Consulting

📦 Logistics

Platform Engineer III managing API integration platforms at Availity. Oversee tooling and support while ensuring flexibility and reliability in healthcare systems integration.

Ansible

AWS

EC2

Splunk

Terraform

🔥 1 hour ago

Ondo Finance

51 - 200

₿ Crypto

💳 Fintech

💸 Finance

Senior Engineer architecting internal AI systems for Ondo Finance. Building solutions connecting data, tools, and workflows to enhance productivity.

🇺🇸 United States – Remote

💰 Initial Coin Offering - Ondo Finance on 2024-01

⏰ Full Time

🟡 Mid-level

🟠 Senior

🏗️ Platform Engineer

🔥 1 hour ago

LMI

1001 - 5000

📦 Logistics

🏥 Healthcare

🎖️ Defense

Power Platform Developer designing and developing SharePoint solutions for LMI. Supporting modernization and automation efforts in a fast-paced environment with a focus on collaboration and agility.

SQL

.NET

🔥 1 hour ago

Dayforce

5001 - 10000

👥 HR Tech

☁️ SaaS

🤝 B2B

Senior Software Engineer for LightBox's AI Platform, creating infrastructure and tools for AI capabilities. Collaborating on AI model integration and production applications.

AWS

Cloud

Distributed Systems

Docker

Google Cloud Platform

JavaScript

Kubernetes

Microservices

Node.js

Python

TypeScript