Senior Monitoring/SRE Engineer

🔥 6 minutes ago

🏛️ District of Columbia, Washington – Remote

infoinfo

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

👻 Ghost score 12%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Koniag Government Services

Koniag Government Services

1001 - 5000 employees

Founded 1975

🏛️ Government

🎖️ Defense

💼 Consulting

Government • Defense • Consulting

Koniag Government Services is an Alaska Native Corporation (ANC) that provides technical, professional, and operational expertise to the U. S. public sector. KGS supports Defense & Intelligence, Federal Civilian, and Health customers with enterprise solutions, professional services, and operations management, and emphasizes mission-focused outcomes, contracting speed (ANC direct awards), and strategic/technology partnerships. The company positions itself as a mission partner delivering people, technology, and program management to government customers.

📋 Description

• Design, implement, and own enterprise monitoring and observability architecture across infrastructure, applications, and cloud services • Define and maintain service level objectives, service level indicators, and error budgets for mission-critical financial and case-management systems • Lead incident response and root cause analysis for high-severity outages • Drive cross-team remediation and after-action reviews • Build automated alerting, dashboards, and runbooks to reduce mean time to detect and mean time to resolve • Support monitoring and validation for the business disaster continuity and recovery program, including synthetic transaction monitoring and failover verification • Partner with server, cloud, storage, and database engineering teams to instrument systems and integrate telemetry into a unified observability platform • Mentor junior SRE/monitoring engineers • Establish best practices for capacity planning and performance baselining • Support continuous monitoring reporting requirements under FISMA/NIST SP 800-53 with the security team • Present operational health, reliability metrics, and improvement roadmaps to program and government leadership

🎯 Requirements

• Minimum of 8 years of experience in systems monitoring, site reliability engineering, or a related operations discipline • Bachelor's degree in Computer Science, Information Technology, or related field, or equivalent professional experience • Hands-on experience designing and administering enterprise monitoring/observability platforms, including Splunk, SolarWinds, Grafana, Datadog, or similar • Strong experience with cloud-native monitoring in AWS, including CloudWatch and CloudTrail or equivalent • Demonstrated experience leading incident response and root cause analysis for enterprise production environments • Experience with AIOps, auto-remediation/self-healing workflows, or OpenTelemetry • Working knowledge of scripting/automation using Python, PowerShell, or Bash • Strong understanding of ITIL-aligned incident, problem, and availability management practices • Excellent written and verbal communication skills, including experience briefing technical and program leadership • Ability to work collaboratively in a fast-paced environment • Ability to convey complex technical concepts to non-technical stakeholders • Ability to obtain public trust clearance • Experience working in a federal government IT environment is desired • Splunk Certified Architect/Admin, AWS Certified DevOps Engineer, or equivalent monitoring/SRE certification is desired • Experience supporting BDCR/COOP monitoring and DR failover validation for financial systems is desired • Experience with PagerDuty, Opsgenie, or similar alert-management/on-call platforms is desired • Familiarity with FedRAMP continuous monitoring (ConMon) reporting requirements is desired

🏖️ Benefits

• Health, dental and vision insurance • 401K with company matching • Flexible spending accounts • Paid holidays • Three weeks paid time off • Competitive compensation • Extraordinary benefits package

Apply Now

Similar Jobs

🔥 10 hours ago

Livefront

201 - 500

💼 Consulting

🤖 Artificial Intelligence

🤝 B2B

DevOps Engineer building reliable cloud infrastructure, automation, and CI/CD systems for Livefront’s AI-enabled software consultancy. Supporting Fortune 1000 client technology initiatives across major U.S. hubs.

🇺🇸 United States – Remote

💵 $80k - $130k / year

💰 Private equity on 2023-10

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🔥 11 hours ago

Alcumus

1001 - 5000

💼 Consulting

🏥 Healthcare

📦 Logistics

Senior DevOps Engineer leading AWS, Kubernetes, Kafka, and CI/CD infrastructure for Veriforce’s contractor-risk management SaaS platform. Managing DevOps engineers and improving production reliability for global customers.

🔥 13 hours ago

Reality Defender (YC W22)

11 - 50

🤖 Artificial Intelligence

🔐 Security

📱 Media

Senior DevOps Engineer owning AWS and Azure infrastructure for Reality Defender’s deepfake detection platform. Automating deployments, Kubernetes operations, observability, security, and reliability.

🔥 13 hours ago

Zscaler

5001 - 10000

🔒 Cybersecurity

☁️ SaaS

🏢 Enterprise

AI DevOps Engineer automating Zscaler’s AWS, GCP, and Azure cloud platform. Integrating Terraform, CI/CD, Python, and AI tooling to secure cybersecurity infrastructure.

🔥 13 hours ago

SYNCREON

10,000+ employees

🚘 Automotive

📦 Logistics

🚗 Transport

Forward Deployment Engineer deploying and integrating cloud software in customer environments. Troubleshooting platforms and supporting implementations for a recruitment and staffing services provider.