Principal Architect/Technical Lead Manager, Site Reliability Engineering
AviatrixAbout the role
The Aviatrix SRE team is a small but highly skilled global group of Systems Engineers/SREs dedicated to ensuring the reliability, availability, and performance of Aviatrix’s critical systems and services. Our mission is to build and maintain a robust, resilient infrastructure that enables Aviatrix to deliver high-quality services with agility through automation, best practices, and a culture of operational excellence.
About the Role
As an SRE – Principal Architect and Technical Lead Manager, you will lead and manage a small team of SRE’s in designing, implementing, and maintaining highly available, fault-tolerant, and scalable systems. You’ll focus on automation, proactive monitoring, and Infrastructure-as-Code (IaC) to drive efficiency and reliability across our services.
Tech Stack & Responsibilities
- Kubernetes – Manage application lifecycles, automate operational tasks, troubleshoot issues, integrate monitoring and alerting, optimize infrastructure, and ensure reliable operations using custom-built operators and cdk8s.
- Terraform – Implement Infrastructure-as-Code (IaC) to enable rapid provisioning, seamless configuration changes, and efficient scaling.
- Automation & Development – Build and enhance automation tools and frameworks in Golang and Python to streamline operations.
On-Call Rotation
We maintain a structured on-call rotation to ensure 24/7 coverage:
Location & Eligibility
This is a remote role open to candidates located in US ideally located on Eastern or Central Time Zone.
RESPONSIBILITIES:
- Lead and mentor a small team of global Site Reliability Engineers
- Ensure Reliability and Availability: You will ensure uptime for crucial services and systems based on business required SLOs. Minimize service disruptions through proactive monitoring, capacity planning and fault-tolerant design.
- Architecture and System Design: you will design and architect complex, scalable and reliable systems.
- Automation and Efficiency: you will develop and implement automation tools and frameworks to automate routine tasks to reduce human error and to streamline and improve operational processes to increase efficiency.
- Build Observability and Monitoring tools: you will define, build, deploy, maintain, and extend our observability and monitoring tools to enhance system reliability and availability.
- Incident Management and Response: you will maintain an effective on-call rotation to ensure 24/7 coverage. You will respond to incident response procedures to swiftly address and mitigate service disruptions.
- Performance Monitoring and SLIs/SLOs: you will help define and monitor Service level Indicators (SLIs) and Service Level Objectives to set clear expectations for system performance.
- Collaboration: you will work closely with product engineering to ensure service-level objectives and reliability targets are met
- Problem-Solving & Troubleshooting: you respond to escalations by troubleshooting complex system and application incidents, perform root cause analysis, implement necessary corrective actions.
- Thought Leadership and Innovation: Stay up to date with latest industry trends, emerging technologies. Iterate on best practices to increase the quality & velocity of development and deliverables.
QUALIFICATIONS:
- 10+ years of experience in software engineering or site reliability engineering roles.
- 2+ years of hands-on leadership role, managing and mentoring SRE engineering teams.
- Proficiency in Golang and Python development skills
- Extensive experience with cloud platforms (e.g., AWS, Azure, GCP) and cloud-native technologies.
- Infrastructure-as-code (IaC): Deep understanding of Terraform core components (e.g., Terragrunt is a bonus) with real-world experience using Terraform for infrastructure provisioning and management.
- Good knowledge of Kubernetes (e.g., cdk8s and operators are a bonus)
- Solid experience developing Automation tools and frameworks.
- Experience with Logging Solutions (e.g., Loki, Syslog, Elasticsearch, Logstash, Kibana, Filebeat, Fluentbit, etc.)
- Experience with Monitoring and Metrics Solutions (e.g., Prometheus, Grafana, Victoria Metrics)
- Practical experience with Linux system administration
- Experience with Version control system (e.g., Git, GitHub) and code review
- Excellent communication skills are required.
US Pay Range
The US National annual base salary range for this full-time position is $244,000 – $265,000 + annual performance bonus
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s