Site Reliability Engineer (SRE)
Alpha OmegaAbout the role
Alpha Omega Integration LLC is an award-winning Federal IT Solutions provider. Since its inception in September 2016, we have grown from a start-up to a $100m/year business. Alpha Omega’s growth stems from our mission focus: to make the US Government the best in the world. We achieve that via advanced capabilities in the areas of Design & Product Management, DevSecOps & Cloud Engineering, Intelligent Automation, and Cybersecurity.
Our consistent growth has fostered a series of accolades including Inc. 5000 and Washington Technology’s Fast 50 awards for five consecutive years, Virginia Business Best Places to Work ten years in a row, and Maryland Technology Council's 2022 Government Contract of the Year over $50 Million Dollars award, to name a few.
We are seeking passionate federal IT professionals to join our team.
Come support our nation’s government agencies and make a difference!
Why Us?
We have H.E.A.R.T.! Alpha Omega's Core Values – (H) harmony, (E) engagement, (A) accountability, (R) resourcefulness, and (T) tenacity- collectively are an acrostic reminder of the values that guide the work we do.
We foster a culture that recognizes and rewards hard work. Our H.E.A.R.T. program invites colleagues and managers from across the organization to recognize each other for living out our core values. Spotlighted employees enjoy a detailed nomination about their core-values-aligned actions which are then shared with their manager.
Ready to embark on a rewarding, challenging, and fulfilling career in the Federal IT Solutions space?
Come grow with us!
Job Title: Software Reliability Engineer (SRE)
Clearance Required: DHS Public Trust EOD
Work Location: Remote
Alpha Omega is seeking a highly skilled and experienced Software Reliability Engineer (SRE) to join our team and lead the reliability efforts for our large-scale Ruby on Rails applications. As an SRE, you will play a pivotal role in ensuring the availability, performance, and security of our mission-critical applications, which supports our refugee and asylum adjudication process worldwide. You will collaborate with approximately 11 teams of talented software engineers, driving the adoption of best practices, tools, and processes to maintain and enhance the reliability of our systems. If you have a passion for solving complex challenges in distributed systems, a deep understanding of Ruby on Rails, Software Reliability Practices, and a commitment to delivering high-quality software, we invite you to apply and help us scale our applications to the next level.
Responsibilities:
- As a Software Reliability Engineer, you will leverage industry-leading tools and practices to ensure the reliability and performance of our large-scale Ruby on Rails applications. You will implement monitoring and observability solutions like Prometheus, Grafana, and New Relic to provide deep insights into system health and performance. Your role will involve automating infrastructure management with Terraform and Docker, streamlining CI/CD pipelines using Jenkins or GitHub Actions, and ensuring code quality through automated testing with RSpec and Brakeman.
- In this role, you will also focus on enhancing system resilience by introducing chaos engineering practices and implementing security measures using tools like OWASP ZAP and 2FA. You will drive the adoption of Infrastructure as Code (IaC), establish standardized processes across teams, and mentor engineers to ensure consistent, high-quality software delivery. Your expertise will be key in maintaining the reliability and security of our applications as they scale.
- To excel as a Software Reliability Engineer, you will need a deep understanding of Ruby on Rails, including performance optimization and troubleshooting within large-scale distributed systems. Proficiency in cloud infrastructure management, container orchestration (e.g., Kubernetes, Docker), and Infrastructure as Code (IaC) tools like Terraform is essential. Strong problem-solving skills are critical, as you will be diagnosing complex issues and implementing automation to streamline processes.
- In addition, you will need effective communication and collaboration skills to work with a large engineering team, ensuring best practices are understood and adopted. Security awareness, including knowledge of web application security principles, is vital to safeguard our systems. Continuous learning and adaptability are also key, enabling you to stay ahead of industry trends and evolving technologies.
Your Role:
In a large team context, the Software Reliability Engineer will play a crucial role in establishing and enforcing reliabil
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s