DevOps Engineer II – IT Infrastructure Systems

🔥 51 minutes ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Golden 1 Credit Union

Golden 1 Credit Union

1001 - 5000 employees

Founded 1933

🏦 Banking

💸 Finance

Banking • Finance

Golden 1 Credit Union is a member-owned credit union headquartered in Sacramento, California. It provides retail and business banking and a wide range of consumer financial services, including multiple checking and savings options, term certificates and IRAs, mortgages and home equity products, auto and personal loans, credit cards, investing and wealth services, and insurance and protection options. The credit union emphasizes online and mobile banking tools, member rewards and cashback programs, financial education, community grants and scholarships, and partnerships (for example with the Golden 1 Center/Sacramento Kings).

📋 Description

• Independently lead infrastructure-as-code development using Terraform and scripting languages such as Python and PowerShell to support scalable and reliable deployments. • Manage Linux/Kubernetes cluster environments. • Deploy solutions in accordance with Change Management Processes. • Support development teams on API integration strategy and standards development. • Ensure systems are secure against cybersecurity threats. • Identify technical problems and develop software updates and fixes. • Strong Splunk skills for administration, query optimization, alerting, and dashboard development. • Build tools to reduce errors and improve customer experience. • Propose ideas and solutions within the Infrastructure Department to reduce workload through automation. • Design, implement, and optimize CI/CD pipelines for faster and more reliable software releases. • Independently conduct root cause analysis and implement corrective actions. • Design and write tests to investigate infrastructure failure and scaling. • Create and maintain response playbooks across incident management and monitoring tools. • Develop automation to ensure repeatability, eliminate toil, and reduce time to action and repair services. • Analyze key operational metrics to identify opportunities to improve availability. • Implement effective monitoring, alerting, and reduction of alert fatigue. • Manage container orchestration environments and optimize deployment workflows to enhance scalability, reliability, and operational efficiency. • Design, build, and manage containerized environments using Docker. • Create and maintain SLIs, SLOs, and error budgets. • Design and optimize monitoring dashboards and alerting systems to proactively detect and address application performance and uptime issues. • Implement code branching strategies using GitHub functions. • Advanced Terraform syntax and GitLab CI/CD configuration, pipelines, jobs. • Provisioning and setting up metrics in Prometheus, Thanos, and Grafana, creating and managing alerts. • Implement cloud engineering standards, reusable modules, and platform patterns in Microsoft Azure. • Operate shared cloud platform services according to Cloud Engineering defined architectures. • Ensure infrastructure changes comply with reliability, security, and cost controls established by Cloud Engineering. • Maintain operational documentation and runbooks for cloud platform services.

🎯 Requirements

• Over 4 years as a DevOps Engineer in medium to large-scale environments. • Proficient in Windows Server, Linux, and hybrid cloud deployments using Microsoft Azure and VMWare. • Skilled in Git/GitHub workflows, Terraform, Python, PowerShell, and container orchestration (Tanzu, Docker, Kubernetes, OpenShift). • Experienced with CI/CD tools (Jenkins, GitLab CI, Azure DevOps) and observability platforms (Datadog, Prometheus, Grafana, ThousandEyes). • Knowledgeable in log management (ELK Stack) and database technologies (PostgreSQL, MySQL, NoSQL). • Strong background in automating infrastructure provisioning and application deployment using Terraform, Ansible, and Kubernetes. • Proficient in creating and maintaining monitoring dashboards, SLIs, SLOs, and error budgets to ensure application uptime and performance. • Experienced in ensuring infrastructure security, driving automation initiatives, and collaborating across teams to improve reliability and scalability. • Experienced in building observability pipelines and performing advanced queries in log management tools like Splunk for troubleshooting. • Experience implementing and operating Azure-based shared services defined by platform or cloud engineering teams. • Microsoft Azure DevOps Engineer Expert Certification (Required). • Kubernetes Administration Certification (Required). • Linux Certification (Desired).

Apply Now

Similar Jobs

🔥 2 hours ago

Counterpart Health

51 - 200

🏥 Healthcare

🤖 Artificial Intelligence

☁️ SaaS

Senior Site Reliability Engineer at Counterpart Health improving infrastructure for healthcare delivery. Collaborating with teams on containerized solutions and automation tools.

AWS

Azure

Cloud

DNS

Docker

Firewalls

Google Cloud Platform

GRPC

Kubernetes

Linux

Prometheus

Python

Shell Scripting

TCP/IP

Go

🔥 3 hours ago

Cognativ

11 - 50

💼 Consulting

🥽 AR/VR

🤖 Artificial Intelligence

Senior Site Reliability Engineer ensuring reliability of a distributed AI video monitoring platform. Leading incident response and managing service reliability and operational quality.

Apache

AWS

Cloud

Grafana

IoT

Java

Kafka

Linux

Postgres

Prometheus

Python

Terraform

Go

🔥 4 hours ago

National Trust

10,000+ employees

🤲 Charity

🏨 Hospitality

🛒 Retail

DevSecOps Engineer III at National Digital Trust Company securing digital asset infrastructure. Lead modernization in CI/CD, security controls, and cloud-native systems.

Cloud

ITSM

Kubernetes

SDLC

Terraform

🔥 4 hours ago

PayNearMe

201 - 500

💳 Fintech

☁️ SaaS

🤝 B2B

Site Reliability Engineer at PayNearMe, Inc. responsible for infrastructure management and application reliability. Collaborating with cross-functional teams to enhance system performance and incident response.

🇺🇸 United States – Remote

💵 $180k - $200k / year

🔥 Funding within the last year

💰 $50M Series E - PayNearMe on 2025-09

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

Ansible

AWS

Azure

Chef

Cloud

Docker

EC2

Google Cloud Platform

Grafana

Kubernetes

Prometheus

Puppet

Python

Ruby

Ruby on Rails

Splunk

Terraform

Go

🔥 4 hours ago

Made4net

51 - 200

📦 Logistics

☁️ SaaS

🏢 Enterprise

Cloud Operations Engineer supporting AWS infrastructure for supply chain software solutions. Monitoring systems and ensuring reliability in a global operations team.

Ansible

AWS

Cloud

DNS

EC2

Grafana

ITSM

Linux

Microservices

Oracle

Postgres

Python

Terraform