Engineer II, Site Reliability

🔥 0 minutes ago

🇬🇧 United Kingdom – Remote

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🇬🇧 UK Skilled Worker Visa Sponsor

info
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of CrowdStrike

CrowdStrike

5001 - 10000 employees

Founded 2011

🔒 Cybersecurity

☁️ SaaS

🤖 Artificial Intelligence

Cybersecurity • SaaS • Artificial Intelligence

CrowdStrike is a cybersecurity company that provides cloud-based security services to stop breaches. It is recognized as a leader in endpoint protection, identity and cloud security, and managed detection and response. CrowdStrike's platform, Falcon, integrates artificial intelligence to offer real-time visibility, detection, and protection against sophisticated cyber threats. The company is lauded for its effectiveness in securing networks and data, making it a trusted partner for businesses worldwide.

📋 Description

• Develop automation and tooling through software for mission-critical solutions and services supporting large-scale distributed systems • Administer and engineer Linux systems across thousands of bare-metal servers and virtual machines • Own platform availability, latency, throughput, monitoring, issue response, deployment, and capacity planning • Participate in an on-call rotation • Troubleshoot server hardware issues • Ensure the platform operates reliably 24x7 • Learn and champion new technologies and techniques across the team • Gain broad exposure to the overall architecture and process flow • Deliver small development projects and occasional larger projects • Use monitoring and telemetry stacks including ELK, Prometheus, Grafana, and Zabbix • Gather and analyze operating-system and application metrics for performance tuning and fault finding • Lead incident analysis, champion incident-response practices, correlate incidents to systemic problems, and drive resolution • Collaborate with globally distributed SREs and engineers • Communicate and present reliability-team conventions • Use AI technologies to enhance decision-making, streamline workflows and processes, improve efficiency, and drive business outcomes

🎯 Requirements

• Bachelor's degree and/or equivalent experience in Computer Science • A minimum of five years of experience working in a large scale production environment • A minimum of two years of experience in software engineering • A minimum of two years of experience in one or more of: C++, Java, Python, Go • Experience with storage technologies such as SAN, NAS, NFS, Object Storage, FreeNAS, and iSCSI • Experience with infrastructure technologies such as Linux, Windows, VMware, Docker, and Kubernetes • Experience writing technical documentation • Configuration management experience with Puppet, Chef, Ansible, or similar tools • Solid understanding of application design and operational trade-offs • Analytical skills coupled with a strong sense of urgency, ownership, and drive • Ability to work well in a diverse, team-focused environment with SREs and Engineers • Ability to broadly communicate and present recommended conventions • Proven experience utilizing AI technologies to enhance decision-making, streamline workflows and processes, improve efficiency, and drive business outcomes

🏖️ Benefits

• Market leader in compensation and equity awards • Comprehensive physical and mental wellness programs • Competitive vacation and holidays for recharge • Paid parental and adoption leaves • Professional development opportunities for all employees regardless of level or role • Employee Networks, geographic neighborhood groups, and volunteer opportunities to build connections • Vibrant office culture with world class amenities • Great Place to Work Certified™ across the globe

Apply Now

Similar Jobs

🕒 3 days ago

Salve.Inno

11 - 50

💼 Consulting

📣 Marketing

📦 Logistics

Senior Site Reliability Engineer operating AWS Kubernetes platforms and improving observability, automation, and incident response. Supporting Salve.Inno Consulting’s clients with scalable, highly available cloud infrastructure.

🕒 3 days ago

Salve.Inno

11 - 50

💼 Consulting

📣 Marketing

📦 Logistics

Senior Site Reliability Engineer operating highly available cloud platforms for Salve.Inno Consulting’s customers worldwide. Automating infrastructure, observability, incident response, and production reliability.

🕒 August 6

PostHog

11 - 50

💼 Consulting

📣 Marketing

📦 Logistics

Site Reliability Engineer scaling PostHog’s petabyte-scale ClickHouse analytics platform on AWS. Automating infrastructure, improving reliability, and reducing operational load for a fast-growing product analytics company.

🕒 August 5

NICE

5001 - 10000

☁️ SaaS

🤖 Artificial Intelligence

📡 Telecommunications

Forward Deployed Engineer building AI agents and full-stack automation for NICE’s customer-experience software. Integrating conversational systems with enterprise platforms and deploying customer self-service solutions.

🕒 August 4

Pliant

201 - 500

💳 Fintech

☁️ SaaS

🤝 B2B

Engineering Manager building Pliant’s Site Reliability function for its B2B payments platform. Establishing SLOs, incident processes, observability, and a new reliability engineering team.