Senior AWS Site Reliability Engineer, SRE

Job not on LinkedIn

🔥 34 minutes ago

🏛️ District of Columbia, Washington – Remote

infoinfo

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

👻 Ghost score 10%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Koniag Government Services

Koniag Government Services

1001 - 5000 employees

Founded 1975

🏛️ Government

🎖️ Defense

💼 Consulting

Government • Defense • Consulting

Koniag Government Services is an Alaska Native Corporation (ANC) that provides technical, professional, and operational expertise to the U. S. public sector. KGS supports Defense & Intelligence, Federal Civilian, and Health customers with enterprise solutions, professional services, and operations management, and emphasizes mission-focused outcomes, contracting speed (ANC direct awards), and strategic/technology partnerships. The company positions itself as a mission partner delivering people, technology, and program management to government customers.

📋 Description

• Design, configure, develop, integrate, test, document, and sustain AWS capabilities • Implement and improve CI/CD pipelines, infrastructure as code, container platform operations, monitoring, alerting, and secure deployment automation • Translate business, mission, security, accessibility, and operational requirements into practical technical solutions • Support platform architecture, backlog refinement, implementation planning, release readiness, and production transition activities • Develop reusable patterns, configuration standards, automation, documentation, and support procedures • Troubleshoot complex issues across platform configuration, code, data, APIs, identity, security, performance, and user experience • Collaborate with cybersecurity, privacy, data, infrastructure, QA, and change management teams • Maintain technical documentation, design decisions, implementation notes, test evidence, and operational runbooks • Provide site reliability, production operations, and cloud platform engineering support for AWS capabilities

🎯 Requirements

• Bachelor's degree in Computer Science, Information Systems, Software Engineering, Data Analytics, Cybersecurity, or a related discipline, or equivalent work experience • 7+ years of experience in site reliability, production operations, and cloud platform engineering • Hands-on experience with AWS implementation, configuration, development, integration, testing, or operations • Experience working with Agile delivery teams and translating stakeholder needs into maintainable technical outcomes • Strong hands-on knowledge of AWS capabilities, implementation patterns, administration, development, integration, and lifecycle management • Ability to design and implement secure, supportable, upgrade-aware solutions • Experience with APIs, identity and access controls, data management, testing, monitoring, troubleshooting, and release coordination • Ability to document technical designs, configuration decisions, operational procedures, test results, and risks • Ability to obtain a Public Trust clearance • Experience with DevSecOps practices, CI/CD pipelines, automated testing, infrastructure as code, or platform release automation • Knowledge of NIST controls, FISMA, FedRAMP-authorized services, audit evidence, and least privilege • Experience improving platform governance, reuse, documentation, observability, and operational readiness • Familiarity with Microsoft 365, ServiceNow, Atlassian, Snowflake, data catalog, CRM, or enterprise integration ecosystems

🏖️ Benefits

• Health, dental and vision insurance • 401K with company matching • Flexible spending accounts • Paid holidays • Three weeks paid time off

Apply Now

Similar Jobs

🔥 5 hours ago

Intellum

51 - 200

💼 Consulting

📣 Marketing

🏥 Healthcare

Lead Site Reliability Engineer modernizing cloud, Kubernetes, and deployment infrastructure for Intellum’s corporate education technology platform. Driving reliability, observability, incident response, and Systems Engineering mentorship.

🔥 7 hours ago

The Home Depot

10,000+ employees

🏗️ Construction

📦 Logistics

🛒 Retail

Senior software engineer building monitoring, orchestration, and observability automation for The Home Depot’s retail and supply chain operations. Modernizing batch workflows and mentoring automation engineers.

🔥 10 hours ago

CDW

10,000+ employees

💼 Consulting

🏥 Healthcare

📚 Education

Senior SRE improving reliability, automation, and observability for CDW’s managed technology services. Troubleshooting production platforms and leading incident response across customer-connectivity environments.

🔥 10 hours ago

Fuze Health

1001 - 5000

🏥 Healthcare

☁️ SaaS

💊 Pharmaceuticals

Senior DevSecOps Engineer securing AWS/GCP infrastructure, Kubernetes, and CI/CD for Fuze Health’s national pharmacy platform. Driving compliance, resilience, and secure engineering at scale.

🔥 11 hours ago

Merative

1001 - 5000

🏥 Healthcare

💼 Consulting

⚕️ Healthcare Insurance

Principal SRE directing reliability, observability, and automation for Merative’s medical imaging cloud platforms. Establishing SLOs, resilience, incident management, and infrastructure automation across teams.