Principal DevOps Engineer

Job not on LinkedIn

🔥 0 minutes ago

🇮🇳 India – Remote

⏰ Full Time

🔴 Lead

⛑ DevOps & Site Reliability Engineer (SRE)

👻 Ghost score 11%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Learning Technologies Group plc

Learning Technologies Group plc

5001 - 10000 employees

Founded 2013

📚 Education

🏢 Enterprise

☁️ SaaS

Education • Enterprise • SaaS

Learning Technologies Group plc (LTG) is a market leader in the fast-growing workplace digital learning and talent management market. LTG offers large organisations a new approach to learning and talent in a business world driven by digital transformation. The company is uniquely positioned to capture growth opportunities in a substantial addressable market, providing solutions for recruitment, motivation, and retention of talent across global companies and government sectors.

📋 Description

• Own and maintain the PeopleFluent Hosting pipeline and automation infrastructure • Own the technical vision and architecture for hosting infrastructure across compute, networking, and platform services • Set technical direction to align design decisions with long-term scalability, security, and maintainability goals • Proactively identify infrastructure optimization opportunities • Work hands-on to scale and secure the company's infrastructure • Lead advanced projects end to end • Document and govern changes through structured processes • Collaborate across engineering and product teams to define requirements • Communicate technical design trade-offs • Mentor others across hosting

🎯 Requirements

• Bachelor's Degree in Computer Science, Engineering, or related field • 10+ years of industry experience • Experience with Enterprise Linux, including SELinux policies, standard configuration and enforcement, and headless server management • Strong network fundamentals, including DHCP and subnetting • Experience with virtualization platforms such as Proxmox and vSphere, including advanced cluster configuration, VLANs, virtual switches, and network segmentation • Experience owning and administering monitoring/observability tools such as Prometheus, Grafana, and Datadog • Experience with Git and shell scripting • Strong knowledge and experience with AWS, including EC2, S3, RDS, and Route53 • Hands-on experience with provisioning, orchestration, and administration using Terraform, Ansible, and Kubernetes • Experience in release and deployment automation using Jenkins • Experience managing configuration changes through a structured change control process, including Git-based workflows, branching, pull requests, and code review • Bonus: Experience in datacenter operations, including automated network switch configuration • Bonus: Experience automating basic firewall rules • Bonus: Database administration with Oracle, MySQL, PostgreSQL, or Microsoft SQL • Bonus: Experience with Elasticsearch and Kafka • Bonus: Experience working in an Agile and Scrum environment • Bonus: Web-based/SaaS company background • Bonus: Startup experience

Apply Now

Similar Jobs

🕒 July 28

Moniepoint Inc. (Formerly TeamApt Inc.)

1001 - 5000

💳 Fintech

🏦 Banking

Engineering Manager leading Site Reliability Engineering tooling delivery at Moniepoint, Africa’s all-in-one financial platform. Driving technical planning, observability, team velocity, and reliable product execution.

JavaScript

Node.js

Python

🕒 June 1

OpenAI

201 - 500

🤖 Artificial Intelligence

☁️ SaaS

🏢 Enterprise

Partner AI Deployment Engineer responsible for AWS deployment strategies and technical leadership in OpenAI. Guiding enterprise customers from ideation to production while influencing joint account strategy.

AWS

🕒 May 18

Akamai Technologies

5001 - 10000

🔒 Cybersecurity

🏢 Enterprise

📱 Media

Site Reliability Engineering Manager leading APJ-based site reliability engineers. Collaborating to define and improve Compute products operation and customer supportability.

Ansible

Chef

Distributed Systems

Puppet

React

SaltStack

🕒 May 13

AlphaSense

1001 - 5000

💼 Consulting

🏥 Healthcare

📣 Marketing

Staff Site Reliability Engineer at AlphaSense enhancing reliability, performance, and scalability of systems. Leading SRE practices and mentoring engineers in a global team.

AWS

Azure

Cloud

DNS

Google Cloud Platform

Grafana

Kubernetes

Prometheus

Python

TCP/IP

Go

🕒 May 13

AlphaSense

1001 - 5000

💼 Consulting

🏥 Healthcare

📣 Marketing

Staff Site Reliability Engineer shaping reliability and performance standards at AlphaSense, driving cultural adoption of SRE best practices across the engineering organization.

AWS

Azure

Cloud

DNS

Google Cloud Platform

Grafana

Kubernetes

Prometheus

Python

TCP/IP

Go