Senior DevOps Engineer, Infrastructure – Reliability

Job not on LinkedIn

🔥 17 hours ago

🐊 Florida – Remote

infoinfo

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

👻 Ghost score 12%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Worth AI

Worth AI

11 - 50 employees

💼 Consulting

🛡️ Insurance

🤖 Artificial Intelligence

Consulting • Insurance • Artificial Intelligence

Worth AI is a company that specializes in enhancing financial security and risk management through advanced AI-driven solutions. Their platform offers a suite of tools designed for seamless onboarding, compliance, and credit risk assessment. They help banks, fintech companies, and credit unions streamline operations by automating processes such as KYC/KYB compliance, reputation monitoring, and automated credit underwriting. Worth AI's technology provides real-time risk monitoring, predictive analytics, and AI-generated insights to improve decision-making, safeguard institutions, and boost revenue. Key offerings include Worth Score™, an AI-driven credit score solution, and unique AI underwriting systems that enhance the accuracy and efficiency of financial assessments.

📋 Description

• Implement scalable Infrastructure-as-Code patterns using Terraform • Own and evolve the Kubernetes platform (EKS or self-managed), ensuring workloads are secure, scalable, and resilient • Optimize CI/CD pipelines to improve deployment frequency, reduce lead time, and increase release confidence • Design and enforce secure networking, IAM, and secrets management strategies across environments • Improve observability through metrics, logs, and tracing using tools such as DataDog • Optimize cloud cost efficiency through rightsizing, autoscaling, and architectural improvements • Implement disaster recovery planning, backup strategies, and multi-region resilience initiatives • Refactor brittle or manually managed infrastructure into automated, testable, reproducible systems • Introduce infrastructure tooling or architectural shifts and drive adoption through documentation, workshops, and hands-on support • Partner with engineering teams to eliminate friction in CI/CD, deployments, and cloud environments • Communicate technical trade-offs across engineering and product stakeholders • Maintain or exceed SLO/SLA targets, reduce incident frequency and duration, improve infrastructure stability and automation, and optimize cloud costs

🎯 Requirements

• 8+ years in DevOps, SRE, or infrastructure engineering • Proven experience designing and operating production Kubernetes environments at scale • Deep hands-on expertise with AWS infrastructure and cloud networking • Strong experience building and maintaining Terraform modules across large cloud environments • Demonstrated ownership of CI/CD systems and measurable improvement of DORA metrics • Experience leading incident response processes and driving meaningful postmortem outcomes • Strong understanding of distributed systems, event-driven architectures (Kafka), and database performance (PostgreSQL) • Proven ability to modernize legacy infrastructure and eliminate manual operational toil • Track record of taking a scoped infrastructure project from an ambiguous starting point to production without needing daily direction • Demonstrated ability to build trust across teams while raising the reliability bar • Bonus: Experience coding applications • Bonus: Experience operating high-throughput Kafka clusters (MSK or self-managed) • Bonus: Strong background in database performance tuning (PostgreSQL, Redis) • Bonus: Experience implementing autoscaling strategies for high-traffic systems • Bonus: Familiarity with service mesh technologies • Bonus: Experience building internal developer platforms (IDP) • Bonus: Background in security best practices (zero-trust networking, policy-as-code) • Bonus: Experience with multi-region or globally distributed systems • Bonus: Experience introducing platform-wide reliability frameworks (SLOs, error budgets, chaos testing)

🏖️ Benefits

• Health Care Plan (Medical, Dental & Vision) • Retirement Plan (401k) • Life Insurance • Flexible Paid Time Off • 9 paid Holidays • Family Leave • Remote • Hybrid work (for Orlando Associates) • Free Food & Snacks (Orlando) • Wellness Resources

Apply Now

Similar Jobs

🔥 19 hours ago

Stratus

501 - 1000

🛡️ Insurance

💼 Consulting

📦 Logistics

Senior SRE operationalizing reliability for Stratus, whose software digitizes MEP contractor workflows. Building SLOs, observability, incident response, and performance engineering across Azure-based production systems.

🔥 22 hours ago

Fleetio

51 - 200

📦 Logistics

💼 Consulting

🚗 Transport

Senior Site Reliability Engineer scaling Fleetio’s Ruby on Rails fleet-management platform. Improving infrastructure reliability, database performance, observability, and AI-powered operations.

🔥 23 hours ago

Scribe

51 - 200

☁️ SaaS

⚡ Productivity

🏢 Enterprise

Senior DevOps Engineer scaling AWS, Kubernetes, and deployment systems for Scribe’s workflow intelligence platform. Ensuring reliability, observability, and cost-efficient infrastructure.

🔥 23 hours ago

CACI International Inc

10,000+ employees

🎖️ Defense

🏛️ Government

🔒 Cybersecurity

Operating System Deployment Engineer managing secure Windows images and deployment automation for CACI’s DoD enterprise IT services. Supporting physical and virtual systems across 187 bases.

🇺🇸 United States – Remote

💵 $63.3k - $129.7k / year

🔥 Funding within the last year

💰 $500M Post-IPO Debt on 2026-02

⏰ Full Time

🟠 Senior

🔴 Lead

⛑ DevOps & Site Reliability Engineer (SRE)

🕒 Yesterday

Guidehouse

10,000+ employees

🏥 Healthcare

🎖️ Defense

📦 Logistics

DevOps Engineer automating cloud infrastructure, CI/CD pipelines, and container deployments for Guidehouse government applications. Supporting secure, reliable software delivery across development, QA, and operations.