Senior Platform/Infrastructure Engineer

Job not on LinkedIn

🔥 0 minutes ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Resilient Co.

Resilient Co.

11 - 50 employees

Founded 2020

🤝 B2B

👥 HR Tech

☁️ SaaS

B2B • HR Tech • SaaS

Resilient Co. is a technology consulting and recruiting firm focused on supporting the growth of startups and organizations within the IT market. With a diverse and experienced team, the company offers a range of services including talent acquisition, coaching, and training tailored to the needs of tech clients. Operating with a 'Remote First' culture, Resilient Co. emphasizes collaboration and adaptability to navigate the challenges of the technology recruitment landscape, positioning itself as a strategic partner for businesses looking to build strong tech teams.

📋 Description

• Design, deploy, and maintain production Kubernetes clusters and related services. • Build and maintain automation and tooling using Python to support platform operations. • Integrate and operate Prometheus for monitoring, alerting, and observability. • Deploy and manage Ceph storage solutions for distributed workloads. • Support platform modernization initiatives and migrate services to cloud-native patterns. • Troubleshoot and resolve issues in distributed systems across compute, storage, and network layers. • Collaborate with development, SRE, and operations teams to define platform requirements and SLAs. • Document platform designs, runbooks, and operational procedures. • Participate in on-call rotations and incident response to maintain platform availability.

🎯 Requirements

• 5+ years of experience in platform, infrastructure, or site reliability engineering roles. • Proven experience deploying and operating Kubernetes in production. • Strong Python skills for automation, tooling, and operational scripts. • Experience implementing and operating Prometheus-based monitoring and alerting. • Hands-on experience with Ceph or similar distributed storage systems. • Cloud experience with AWS and Azure (designing, deploying, and operating services). • Demonstrated ability to troubleshoot distributed systems and resolve production incidents. • Experience collaborating across teams to deliver platform improvements and migrations. • Experience with OpenSearch (nice to have). • Proficiency with Bash scripting (nice to have). • Familiarity with Java-based services (nice to have). • Experience with Fluent Bit for log collection (nice to have). • Experience working with PostgreSQL (nice to have).

🏖️ Benefits

• Laptop: BYOD. • Overtime Required: No. • Client Holidays (USA – Mandatory)

Apply Now

Similar Jobs

🕒 July 11

HomeVision

11 - 50

💳 Fintech

🤖 Artificial Intelligence

🏢 Enterprise

Infrastructure Engineer focusing on AI and data security for software in the mortgage industry. Working on secure infrastructure and data-access security while managing independent projects.

AWS

Python

Terraform

TypeScript

Go

🕒 June 24

Nexus

11 - 50

🤖 Artificial Intelligence

🔌 API

🔒 Cybersecurity

Infrastructure Engineer supporting the core infrastructure of Nexus's Layer 1 blockchain. Collaborate remotely in LATAM to enhance the high-performance exchange operations.

AWS

Azure

Cloud

Distributed Systems

Docker

Google Cloud Platform

Grafana

Kubernetes

Prometheus

Rust

Terraform

Go