Senior Monitoring/SRE Engineer (REMOTE)
Koniag Government ServicesAbout the role
Koniag IT Systems, LLC, a Koniag Government Services company, is seeking a Senior Monitoring/SRE Engineer to support KITS and our government customer in Washington, DC. The position is remote. This position requires the candidate to be able to obtain a Public Trust.
We offer competitive compensation and an extraordinary benefits package including health, dental and vision insurance, 401K with company matching, flexible spending accounts, paid holidays, three weeks paid time off, and more.
Koniag IT Systems (KITS) is seeking a Senior Monitoring / SRE Engineer with a minimum of 8 years of experience to lead enterprise observability, monitoring, and reliability engineering practices for a federal civilian customer's hybrid on-premises and AWS cloud environment. The ideal candidate has designed and operated enterprise monitoring platforms at scale, has strong incident management experience, and can drive site reliability practices across a large, multi-team infrastructure program.
The Senior Monitoring / SRE Engineer will be responsible for designing, implementing, and owning the enterprise monitoring and observability architecture, leading incident response and root cause analysis for high-severity outages, and supporting continuous monitoring (ConMon) reporting requirements under FISMA/NIST SP 800-53.
Key responsibilities include but are not limited to:
• Design, implement, and own the enterprise monitoring and observability architecture spanning infrastructure, applications, and cloud services (e.g., Splunk, SolarWinds, Grafana, Datadog, or CloudWatch).
• Define and maintain service level objectives (SLOs), service level indicators (SLIs), and error budgets for mission-critical financial and case-management systems.
• Lead incident response and root cause analysis for high-severity outages, driving cross-team remediation and after-action reviews.
• Build automated alerting, dashboards, and runbooks to reduce mean time to detect (MTTD) and mean time to resolve (MTTR).
• Support monitoring and validation activities for the customer's business disaster continuity and recovery (BDCR) program, including synthetic transaction monitoring and failover verification.
• Partner with server, cloud, storage, and database engineering teams to instrument systems and integrate telemetry into a unified observability platform.
• Mentor junior SRE/monitoring engineers and establish best practices for capacity planning and performance baselining.
• Support continuous monitoring (ConMon) reporting requirements under FISMA/NIST SP 800-53 in coordination with the security team.
• Present operational health, reliability metrics, and improvement roadmaps to program and government leadership.
Education and Experience:
Required:
• Bachelor's degree in Computer Science, Information Technology, or related field, or equivalent professional experience
• Minimum of 8 years of experience in systems monitoring, site reliability engineering, or a related operations discipline
• Hands-on experience designing and administering enterprise monitoring/observability platforms (Splunk, SolarWinds, Grafana, Datadog, or similar)
• Strong experience with cloud-native monitoring in AWS (CloudWatch, CloudTrail, or equivalent)
• Demonstrated experience leading incident response and root cause analysis for enterprise production environments
• Experience with AIOps, auto-remediation/self-healing workflows or OpenTelemetry
Required Skills and Competencies:
• Working knowledge of scripting/automation (Python, PowerShell, or Bash) for monitoring and alerting integration
• Strong understanding of ITIL-aligned incident, problem, and availability management practices
• Excellent written and verbal communication skills, including experience briefing technical and program leadership
• Ability to work collaboratively in a fast-paced environment
• Excellent communication skills and the ability to convey complex technical concepts to non-technical stakeholders
• Ability to obtain public trust clearance
Desired Skills and Competencies:
• Experience working in a federal government IT environment
• Splunk Certified Architect/Admin, AWS Certified DevOps Engineer, or equivalent monitoring/SRE certification
• Experience supporting BDCR/COOP monitoring and DR failover validation for financial systems
• Experience with PagerDuty, Opsgenie, or similar alert-management/on-call platforms
• Familiarity with FedRAMP continuous monitoring (ConMon) reporting requirements
Our Equal Employment Opportunity Policy
The company is an equal opportunity employer. The company shall not discriminate against any employee or applicant because of race, color, religion, creed, ethnicity, sex, sexual orientation, gender or gender identity (
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s