Senior Site Reliability Engineer
Eagle Eye NetworksAbout the role
About Us
Eagle Eye Networks is the global leader in cloud video surveillance, delivering cyber-secure, cloud-based video with artificial intelligence (AI) and analytics to make businesses more efficient and the world a safer place. The Eagle Eye Cloud VMS (video management system) is the only platform robust and flexible enough to power the future of video surveillance and intelligence. Eagle Eye is based in Austin, Texas, with offices in Amsterdam, Bangalore, and Tokyo. Learn more at een.com.
Summary
Eagle Eye Networks is looking for an experienced Senior Site Reliability Engineer (SRE) to improve reliability, performance, and engineering excellence across our global video surveillance platform. In this role, you will operate and automate infrastructure, lead incident response, enhance on-call processes, drive reliability and observability initiatives, and mentor engineers. You will connect development and operations, taking ownership of outcomes and raising the standards for resilient systems. If you are a technical leader who can identify systemic issues, manage outages, and mentor others, you will thrive here.
Responsibilities
- Design and maintain resilient, automated infrastructure in private cloud environments.
- Lead incident response efforts, including communication and follow-ups during major incidents.
- Drive initiatives to reduce recurring issues and enhance both availability and recovery.
- Define and enforce best practices for observability, incident management, and production readiness to ensure optimal performance and reliability.
- Lead improvements in Infrastructure as Code, CI/CD tooling.
- Collaborate with product and application teams to define Service Level Indicators (SLIs) and Service Level Objectives (SLOs) while addressing reliability risks.
- Advocate for automation and operational efficiency.
- Contribute to the reliability roadmap and engage in architecture discussions.
- Mentor engineers and promote a culture of learning and ownership.
- Participate in the on-call rotation and work to improve its effectiveness.
Must Have:
- 5+ years of experience as a Site Reliability Engineer (SRE).
- Demonstrates mastery in managing Linux systems in production environments with ease and is able to effectively teach others.
- Proficient in leveraging Kubernetes and similar container orchestration systems, possessing a level of expertise that matches the depth of my Linux administration skills..
- Demonstrates an advanced proficiency in scripting and programming, particularly with languages such as Python, Bash, and Golang.
- Prior experience mentoring engineers or leading reliability initiatives.
- Experience in building automation tools to reduce operational toil and improve service availability.
- Proven experience leading incident response and conducting root cause analysis.
- Experience with maintaining and using Prometheus(or VictoriaMetrics), Grafana, and other observability tool
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s