Site Reliability Engineer
By Light Professional IT Services LLCAbout the role
Company Overview
By Light Professional IT Services LLC readies warfighters and federal agencies with technology and systems engineered to connect, protect, and prepare individuals and teams for whatever comes next. Headquartered in McLean, VA, By Light supports defense, civilian, and commercial IT customers worldwide.
Cole Engineering Services (CESI), a By Light company, is recognized as a premier provider of modeling and simulation (M&S) training solutions to the Federal Government and industry. Since 2004, CESI has been at the forefront of developing, maintaining, and integrating simulation-based training, serious gaming, technical services, training and other support in live, virtual, constructive, and gaming (LVCG) domains. CESI also designs, builds and runs infrastructure, platforms, applications and processes that enable cyber training for the integrated multi-domain force. Our vision is to become a worldwide full spectrum LVCG and cyber training/analysis developer, integrator and services provider.
Position Overview
Proactively find system deficiencies, prioritize backlogs, and lead the team to correct issues before they impact users. Coordinate with development engineering teams and follow program processes to ensure controlled changes to production systems. Identify and implement automation, observability, and configuration updates throughout the system. Work as part of a fast moving highly diverse team across multiple projects simultaneously.
We have multiple positons available at carying levels and shifts. Shifts include 6am-2pm and 2pm-10pm.
Required Experience/Qualifications
Candidates do not have to be an expert in all but experience in these will be heavily weighted:
- Must have experience leading a technical team of engineers, including performing supervisory duties to ensure successful team outcomes.
- Must be able to understand complex enterprise IT system architectures (hardware and software) with the ability to pinpoint the team’s focus areas and prioritizations to maximize impact.
- Must be a self-starter in a fast-paced environment and able to work with a multi-faceted technical team of engineers with a diverse set of skills at differing levels of experience.
- Establish processes and procedures for maintaining infrastructure configuration management across the enterprise to include automated system checks against the configuration managed baseline.
- Must be able to effectively manage processes associated with the team’s use of Git repositories to ensure proper configuration management.
- Analyzes real world problems and implements solutions according to corporate and government guidelines, procedures, and industry best practices.
- Assures system stability, accessibility, and proper configuration of assigned technical systems and components.Be sensitive and flexible to the needs and requirements of the customer.
- Must be comfortable with Linux system management, as this is the primary operating system, as well as other Red Hat enterprise services.
- Experience with VMware Cloud Foundations, and the overarching VMware hypervisor and management stack and servicesOperate, and maintain physical switches, routers, IPS, IDS devices.
- Operate, and maintain a VMware NSX based SDN.
- Understand Kubernetes and micro-service based applications.
- Must be able to author and maintain scripts and automation in a managed Git repository.
- Proactively identify, locate, mitigate, and resolve system issues.
- Document and present performance records and results.
- Document problems/issues via Jira based ticketing system/tracking log.
- Review and update all information security management system process and procedures (data/software).
- Bachelor’s degree in a technical discipline such as computer science or information technology from an accredited college or university. Will consider additional degrees with accompanied technical leadership experience.
- Ten years of work experience preferred. Security+ certifications are required or must be completed within six months of hire.
Preferred Experience/Qualifications
- Experience with Ansible (or similar) infrastructure automation to deploy and configure standard baselines for physical devices and in a virtual environment (VMware and Linux).
- Utilize scripting for automating tasks, administration, data collection and reporting. This job requires the ability to write, test, debug, deploy, and maintain scripts.
- Ability to audit and report system and performance logs.
- Ability to automate, administer and maintain the environment through configuration management, version control, and backups with the use of scripting.
- Ability to respond professionally, effectively, and efficiently to ser
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s