Senior Platform Engineer

🔥 4 minutes ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Flexential

Flexential

501 - 1000 employees

Founded 2000

🤝 B2B

🏢 Enterprise

📡 Telecommunications

B2B • Enterprise • Telecommunications

Flexential is a US-based provider of data center infrastructure and hybrid IT services. The FlexAnywhere® platform combines colocation, cloud (public/private/hybrid), connectivity/interconnection, data protection, managed and professional services to support enterprise workloads — including high-density GPU and AI/ML deployments. Flexential operates 40+ data centers across 18 US markets (3M+ sq. ft. footprint) and a 100+ Gbps network backbone, and positions itself as a B2B partner for enterprises needing resilient, scalable, and interconnected infrastructure.

📋 Description

• Design, develop and operationally manage automated, resilient, high availability, self-healing, secure platforms with native-AI capabilities for IT needs, serving both internal as well as customer business capabilities. • Develop, and manage the Observability OpenTelemetry Central Backend Stack: Grafana Enterprise, Mimir, Loki, Tempo, and Alertmanager on Kubernetes/RKE2 via Helm and GitLab CI-CD. • Build and manage iaC and CI-CD for automated provisiong and deployment, including Terraform modules for Infra/VM/storage provisioning, Ansible AWX playbooks for OS/App bootstrap, ArgoCD and Helm for Kubernetes configuration. • Develop and manage OpenTelemetry Prometheus scrape profile library including SNMP exporters, REST API exporters, and cloud provider exporters (CloudWatch, Azure Monitor, GCP) for multiple device classes. • Develop AIOps capabilities on platforms for e.g Observability use-cases: anomaly detection integrations, event correlation rules in Alertmanager, and synthetic monitoring patterns to reduce alert noise. • Configure and maintain Zabbix auto-discovery: network range scanning, device classification, and Prometheus service discovery integration. • Build and harden Edge Stack deployments (Prometheus + OTel collector) per data center site using GitOps templates. • Integrate Alertmanager with ServiceNow: webhook routing, ticket enrichment, auto-close logic, and escalation policy configuration. • Maintain platform security: Conjur/CyberArk secret injection at runtime, mTLS between stack components, RBAC in Grafana Enterprise. • Author and maintain Grafana dashboards in JSON/GitLab — facility overview, network health, RED metrics, application telemetry. • Mentor mid-level engineers, lead code reviews, and establish engineering standards for the team. • Represent platform engineering in cross-functional architecture reviews and executive-level program updates. • Perform other duties as required and assigned.

🎯 Requirements

• 5+ years in a production environment • Kubernetes (RKE2/k3s) • Helm chart deployment • System services • Docker/container • 4+ years: Grafana, Mimir, Loki, Tempo configuration, tuning, dash-boarding and production operations • Prometheus required • 5+ years senior-level Python / Scripting frameworks • Automation scripts • Exporter development • GitLab pipeline scripting • REST API integrations • 5+ years GitOps / CI/CD • GitLab CI/CD pipeline authoring • Terraform and Ansible as primary IaC tools • ArgoCD or Flux preferred • 2+ years AIOps / Observability Engineering • Alertmanager rule authoring • Anomaly detection integration • Event correlation • Noise reduction techniques • 5+ years Working Infrastructure (Linux/VM) Management Knowledge • Linux administration • VMware vCenter/VCF experience • Netapp storage management • Network fundamentals (SNMP, TCP/IP) • 2+ years Secrets Management • CyberArk/Conjur, HashiCorp Vault, or equivalent • Runtime secret injection patterns • Minimal travel may be required

🏖️ Benefits

• Medical, Telehealth, Dental and Vision • 401(k) • Health Savings Accounts (HSA) and Flexible Spending Accounts (FSA) • Life and AD&D • Short Term and Long-Term disability • Flex Paid Time Off (PTO) • Leave of Absence • Employee Assistance Program • Wellness Program • Rewards and Recognition Program

Apply Now

Similar Jobs

🔥 59 minutes ago

Quantiphi

1001 - 5000

💼 Consulting

🏥 Healthcare

📦 Logistics

Senior Platform Engineer at Quantiphi designing and optimizing infrastructure for GenAI and LLM workloads. Collaborating with data science and application teams to deliver AI solutions.

🔥 1 hour ago

Availity

1001 - 5000

🏥 Healthcare

💼 Consulting

📦 Logistics

Platform Engineer III managing API integration platforms at Availity. Oversee tooling and support while ensuring flexibility and reliability in healthcare systems integration.

🔥 1 hour ago

Ondo Finance

51 - 200

₿ Crypto

💳 Fintech

💸 Finance

Senior Engineer architecting internal AI systems for Ondo Finance. Building solutions connecting data, tools, and workflows to enhance productivity.

🇺🇸 United States – Remote

💰 Initial Coin Offering - Ondo Finance on 2024-01

⏰ Full Time

🟡 Mid-level

🟠 Senior

🏗️ Platform Engineer

🔥 1 hour ago

LMI

1001 - 5000

📦 Logistics

🏥 Healthcare

🎖️ Defense

Power Platform Developer designing and developing SharePoint solutions for LMI. Supporting modernization and automation efforts in a fast-paced environment with a focus on collaboration and agility.

🔥 1 hour ago

Dayforce

5001 - 10000

👥 HR Tech

☁️ SaaS

🤝 B2B

Senior Software Engineer for LightBox's AI Platform, creating infrastructure and tools for AI capabilities. Collaborating on AI model integration and production applications.