Lead Site Reliability Engineer

🔥 0 minutes ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of IPolarity

IPolarity

51 - 200 employees

Founded 2011

💼 Consulting

🏥 Healthcare

📦 Logistics

Consulting • Healthcare • Logistics

IPolarity is a professional IT services firm that provides cloud computing, virtualization, data center management, CRM implementation, data science, machine learning, and NLP/chatbot solutions to enterprise clients. The company offers both consulting and product development (including EasySAS, Panorama 360, and Orion) and supports workforce strategies such as staff augmentation and hiring, serving B2B customers across industries from offices in the US, Canada, and India.

📋 Description

• Design, deploy, and operate enterprise observability platforms. • Build and maintain Splunk Enterprise/Splunk Cloud infrastructure including Indexers, Search Head Clusters, Heavy Forwarders, and Deployment Servers. • Deploy and operate large-scale Elasticsearch clusters for log analytics and search. • Design, deploy, and support distributed tracing platforms using Grafana Tempo and OpenTelemetry. • Build and maintain end-to-end tracing pipelines, instrumentation standards, and trace retention strategies. • Scale Prometheus, Grafana, Kafka, Tempo, and OpenTelemetry-based monitoring solutions. • Develop dashboards, alerts, analytics, and trace visualizations using Splunk SPL, Grafana, Kibana, and Tempo. • Automate infrastructure using Terraform and configuration management tools.

🎯 Requirements

• 7+ years in Site Reliability Engineering, Platform Engineering, or DevOps. • Hands-on experience administering Splunk Enterprise or Splunk Cloud. • Strong knowledge of Splunk SPL. • Experience with Elasticsearch/ELK, Prometheus, Grafana, Grafana Tempo, distributed tracing, OpenTelemetry, and Kafka. • Experience implementing metrics, logs, and traces as part of a modern observability strategy. • Experience with Terraform and Infrastructure as Code. • Programming experience in Python, Go, Ruby, or Bash. • Splunk certification preferred. • Experience with Kubernetes, AWS/Azure/GCP, Ansible, Consul, CI/CD pipelines, and service mesh technologies.

Apply Now

Similar Jobs

🔥 1 hour ago

Aya Healthcare

5001 - 10000

🏥 Healthcare

💼 Consulting

📦 Logistics

Manager of Site Reliability Engineering leading a team for Aya Healthcare's workforce platform. Ensuring product reliability and outstanding user experience through innovative solutions.

🔥 4 hours ago

Hearst Health

1001 - 5000

🏥 Healthcare

⚕️ Healthcare Insurance

☁️ SaaS

Senior DevOps Engineer II at Bring a Trailer modernizing the infrastructure and security of a trusted automotive marketplace. Developing applications to enable engineering teams to work efficiently.

🔥 7 hours ago

Vynca

51 - 200

💼 Consulting

⚕️ Healthcare Insurance

🏥 Healthcare

Site Reliability Engineer designing and managing scalable AWS infrastructure for Vynca's healthcare tech platform. Collaborating with Software Engineers and improving system reliability through automation and observability practices.

🔥 8 hours ago

Akkadian Labs

51 - 200

☁️ SaaS

🏢 Enterprise

📡 Telecommunications

DevOps Engineer supporting design, implementation, and maintenance of secure infrastructure. Collaborating with teams to enable reliable deployments and improve system observability at Akkadian Labs.

🔥 8 hours ago

Sentara Health

10,000+ employees

🏥 Healthcare

⚕️ Healthcare Insurance

DevOps Engineer focusing on developing Azure cloud infrastructure and Kubernetes environments for Sentara Health. Managing CI/CD pipelines and collaborating with development teams in a remote setting.