Principal Site Reliability Engineer
Clarity InnovationsAbout the role
Clarity Innovations is a trusted national security partner, dedicated to safeguarding our nation’s interests and delivering innovative solutions that empower the Intelligence Community (IC) and Department of Defense (DoD) to transform data into actionable intelligence, ensuring mission success in an evolving world.
Our mission-first software and data engineering platform modernizes data operations, utilizing advanced workflows, CI/CD, and secure DevSecOps practices. We focus on challenges in Information Warfare, Cyber Operations, Operational Security, and Data Structuring, enabling end-to-end solutions that drive operational impact.
We are committed to delivering cutting-edge tools and capabilities that address the most complex national security challenges, empowering our partners to stay ahead of emerging threats and ensuring the success of their critical missions. At Clarity, we are people-focused and set on being a destination employer for top talent, offering an environment where innovation thrives, careers grow, and individuals are valued. Join us as we continue to lead innovation and tackle the most pressing challenges in national security.
Principal Site Reliability Engineer
Description:
We are seeking a Principal Site Reliability Engineer (SRE) to support a mission-critical, on-premises “IT department in a box” model. This role ensures the reliability, security, and performance of a highly complex, compliance-driven hybrid infrastructure. You will provide broad support — from architecture and automation to daily system operations and hardening — with a focus on Red Hat Enterprise Linux (RHEL), Windows integration, and VMware-based virtualization.
As a senior technical leader, you will drive the execution of site reliability engineering across system design, observability, security posture, and automation frameworks. This role requires deep experience across infrastructure layers and a strong commitment to operational excellence, security compliance, and continuous improvement.
Key Responsibilities
Mission-Critical Reliability & Recovery
- Lead response and remediation efforts for complex, high-impact infrastructure issues such as:
- Large-scale file permission repairs
- System and service recovery
- Act as the escalation point for high-severity incidents, providing root cause analysis and durable fixes.
Linux & Hybrid Infrastructure Expertise
- Apply deep knowledge of Red Hat Enterprise Linux (RHEL) at scale, including SELinux configuration and enforcement.
- Design and implement secure, high-availability systems that span VMware, Windows, and Linux environments.
- Manage distributed file operations and migration using tools such as rsync, find, and advanced shell scripting.
Automation and Configuration Management
- Lead the development of infrastructure-as-code using Ansible and YAML for system provisioning, state enforcement, and compliance.
- Leverage Red Hat Satellite or equivalent tooling to manage configuration drift, patching, and lifecycle policies.
Security Posture & Compliance Alignment
- Own the implementation and enforcement of security hardening measures, including file integrity, SELinux policy tuning, access controls, and secure configuration baselines.
- Collaborate with security teams to maintain posture in line with compliance frameworks (e.g., CIS Benchmarks, NIST, STIGs).
Identity Federation & Systems Integration
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s