
1001 - 5000 employees
Founded 1975
🏛️ Government
🎖️ Defense
💼 Consulting
Government • Defense • Consulting
Koniag Government Services is an Alaska Native Corporation (ANC) that provides technical, professional, and operational expertise to the U. S. public sector. KGS supports Defense & Intelligence, Federal Civilian, and Health customers with enterprise solutions, professional services, and operations management, and emphasizes mission-focused outcomes, contracting speed (ANC direct awards), and strategic/technology partnerships. The company positions itself as a mission partner delivering people, technology, and program management to government customers.
🕒 August 21
🏛️ District of Columbia, Washington – Remote
⏰ Full Time
🟠 Senior
⛑ DevOps & Site Reliability Engineer (SRE)
👻 Ghost score 13%
Improve your chances of getting an interview by checking your resume score before you apply.

1001 - 5000 employees
Founded 1975
🏛️ Government
🎖️ Defense
💼 Consulting
Government • Defense • Consulting
Koniag Government Services is an Alaska Native Corporation (ANC) that provides technical, professional, and operational expertise to the U. S. public sector. KGS supports Defense & Intelligence, Federal Civilian, and Health customers with enterprise solutions, professional services, and operations management, and emphasizes mission-focused outcomes, contracting speed (ANC direct awards), and strategic/technology partnerships. The company positions itself as a mission partner delivering people, technology, and program management to government customers.
• Design, implement, and own enterprise monitoring and observability architecture across infrastructure, applications, and cloud services • Define and maintain service level objectives, service level indicators, and error budgets for mission-critical financial and case-management systems • Lead incident response and root cause analysis for high-severity outages • Drive cross-team remediation and after-action reviews • Build automated alerting, dashboards, and runbooks to reduce mean time to detect and mean time to resolve • Support monitoring and validation for the business disaster continuity and recovery program, including synthetic transaction monitoring and failover verification • Partner with server, cloud, storage, and database engineering teams to instrument systems and integrate telemetry into a unified observability platform • Mentor junior SRE/monitoring engineers • Establish best practices for capacity planning and performance baselining • Support continuous monitoring reporting requirements under FISMA/NIST SP 800-53 with the security team • Present operational health, reliability metrics, and improvement roadmaps to program and government leadership
• Minimum of 8 years of experience in systems monitoring, site reliability engineering, or a related operations discipline • Bachelor's degree in Computer Science, Information Technology, or related field, or equivalent professional experience • Hands-on experience designing and administering enterprise monitoring/observability platforms, including Splunk, SolarWinds, Grafana, Datadog, or similar • Strong experience with cloud-native monitoring in AWS, including CloudWatch and CloudTrail or equivalent • Demonstrated experience leading incident response and root cause analysis for enterprise production environments • Experience with AIOps, auto-remediation/self-healing workflows, or OpenTelemetry • Working knowledge of scripting/automation using Python, PowerShell, or Bash • Strong understanding of ITIL-aligned incident, problem, and availability management practices • Excellent written and verbal communication skills, including experience briefing technical and program leadership • Ability to work collaboratively in a fast-paced environment • Ability to convey complex technical concepts to non-technical stakeholders • Ability to obtain public trust clearance • Experience working in a federal government IT environment is desired • Splunk Certified Architect/Admin, AWS Certified DevOps Engineer, or equivalent monitoring/SRE certification is desired • Experience supporting BDCR/COOP monitoring and DR failover validation for financial systems is desired • Experience with PagerDuty, Opsgenie, or similar alert-management/on-call platforms is desired • Familiarity with FedRAMP continuous monitoring (ConMon) reporting requirements is desired
• Health, dental and vision insurance • 401K with company matching • Flexible spending accounts • Paid holidays • Three weeks paid time off • Competitive compensation • Extraordinary benefits package
Apply Now🕒 August 20
AI DevOps Engineer automating Zscaler’s AWS, GCP, and Azure cloud platform. Integrating Terraform, CI/CD, Python, and AI tooling to secure cybersecurity infrastructure.
🇺🇸 United States – Remote
💵 $140k - $175k / year
💰 Secondary Market on 2017-11
⏰ Full Time
🟡 Mid-level
🟠 Senior
⛑ DevOps & Site Reliability Engineer (SRE)
🦅 H1B Visa Sponsor
🕒 August 19
Forward Deployed DevOps Engineer shipping cloud infrastructure for Forge’s AI-native DevOps workspace. Designing automation, improving reliability, and embedding with customer engineering teams.
🕒 August 18
Senior cloud deployment engineer delivering secure Azure, AWS, and hybrid infrastructure for business and government customers. Leading migrations, automation, networking, security, and modernization initiatives.
🇺🇸 United States – Remote
⏰ Full Time
🟠 Senior
⛑ DevOps & Site Reliability Engineer (SRE)
🦅 H1B Visa Sponsor
🕒 August 18
Data center engineer designing POP infrastructure, rack layouts, power systems, and deployment documentation. Supporting Astreya’s global IT managed services through network infrastructure projects and vendor coordination.
🇺🇸 United States – Remote
💵 $73k - $115.2k / year
⏰ Full Time
🟡 Mid-level
🟠 Senior
⛑ DevOps & Site Reliability Engineer (SRE)
🦅 H1B Visa Sponsor
Switching
🕒 August 18
Lead DevOps Engineer owning compliant AWS environments, databases, CI/CD, and observability for legal-industry software. Mentoring platform engineers and enabling secure, reliable product delivery.
🇺🇸 United States – Remote
💵 $165k - $205k / year
⏰ Full Time
🟠 Senior
⛑ DevOps & Site Reliability Engineer (SRE)