Infrastructure – DevOps Lead

Job not on LinkedIn

🕒 July 13

🇬🇧 United Kingdom – Remote

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🇬🇧 UK Skilled Worker Visa Sponsor

infoinfo

👻 Ghost score 22%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Brahma

Brahma

11 - 50 employees

Founded 2022

₿ Crypto

💳 Fintech

🔌 API

💰 $4.2M Seed Round - Brahma on 2022-02

Crypto • Fintech • API

Brahma is a developer-focused orchestration layer that connects onchain financial logic to offchain payment and settlement systems. It enables programmable capital flows across blockchains, strategies, and traditional payment rails, providing smart accounts, automated agents, and programmable card issuance so apps can turn onchain positions into real-world spending. Brahma abstracts infrastructure and settlement for developers, letting them build onchain or hybrid apps without running backend ops.

📋 Description

• Lead, mentor, and grow a team of 7 DevOps and Infrastructure engineers. • Drive agile delivery, sprint planning, and backlog prioritisation to align infrastructure deliverables with AI research and product roadmaps. • Establish best practices for Reliability Engineering, Infrastructure-as-Code (IaC), continuous integration, and incident post-mortems. • Manage high-density GPU clusters across a multi-cloud ecosystem optimised for large custom AI model training and real-time inference workflows. • Oversee infrastructure consumption, track cloud/hardware costs, negotiate vendor terms, and optimise GPU utilisation. • Serve as the senior technical escalation point for complex infrastructure incidents and architecture decisions. • Standardise platform deployments using Infrastructure as Code and modern container orchestration. • Partner with security stakeholders to ensure our AI training environments meet industry security standards.

🎯 Requirements

• Proven track record leading or managing a team of 5+ infrastructure, platform, or DevOps engineers. • Hands-on experience architecting and managing GPU-intensive workloads (NVIDIA clusters, cloud AI accelerators) for compute-heavy applications. • Expertise with Kubernetes, Docker, Terraform (or OpenTofu), and multi-cloud environments (with strong hands-on GCP experience). • Demonstrated experience designing, optimising, and maintaining high-performance storage architectures and caching layers for demanding compute workloads. • Strong experience with cloud cost governance (FinOps), capacity planning, and vendor interaction. • Exceptional stakeholder management skills with the ability to bridge business requirements and deep technical infrastructure details.

🏖️ Benefits

• Health insurance • Professional development opportunities

Apply Now

Similar Jobs

🕒 July 9

Ensono

1001 - 5000

💼 Consulting

Site Reliability Engineer managing Cloud and Infrastructure as Code at Ensono. Leading client-facing discussions and driving service improvement initiatives.

🕒 July 6

Omilia - Conversational Intelligence

201 - 500

💼 Consulting

🛡️ Insurance

✈️ Travel

Senior Site Reliability Engineer operating and maintaining production clusters while developing observability solutions. Collaborating with teams to enhance platform reliability through automation and monitoring.

🕒 July 6

DeepHealth

11 - 50

🏥 Healthcare

💼 Consulting

📦 Logistics

DevOps Engineer managing AWS infrastructure and enhancing platform reliability at DeepHealth. Collaborating with teams to automate processes and improve software delivery.

🇬🇧 United Kingdom – Remote

💵 £60k - £70k / year

💰 $225k Grant on 2019-08

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🕒 July 6

DeepHealth

11 - 50

🏥 Healthcare

💼 Consulting

📦 Logistics

DevOps Engineer responsible for AWS and Kubernetes platform management at DeepHealth, a healthcare SaaS provider. Ensuring cloud infrastructure is secure and reliable for efficient software delivery.

🇬🇧 United Kingdom – Remote

💵 £60k - £70k / year

💰 $225k Grant on 2019-08

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🕒 July 2

RTX

10,000+ employees

🏭 Manufacturing

💼 Consulting

📦 Logistics

Principal Site Reliability Engineer managing complex AWS infrastructures for Collins Aerospace. Delivering B2B solutions and overseeing cloud migration projects in a collaborative team environment.