Jobs and Careers
AV

Principal Architect/Technical Lead Manager, Site Reliability Engineering

Aviatrix
United States, United Statesfull_timeVerifiedPosted 17 Mar 2025
💰 $265,000/yr($244,000/yr$265,000/yr)

About the role

The Aviatrix SRE team is a small but highly skilled global group of Systems Engineers/SREs dedicated to ensuring the reliability, availability, and performance of Aviatrix’s critical systems and services. Our mission is to build and maintain a robust, resilient infrastructure that enables Aviatrix to deliver high-quality services with agility through automation, best practices, and a culture of operational excellence.

About the Role

As an SRE – Principal Architect and Technical Lead Manager, you will lead and manage a small team of SRE’s in designing, implementing, and maintaining highly available, fault-tolerant, and scalable systems. You’ll focus on automation, proactive monitoring, and Infrastructure-as-Code (IaC) to drive efficiency and reliability across our services.

Tech Stack & Responsibilities

  • Kubernetes – Manage application lifecycles, automate operational tasks, troubleshoot issues, integrate monitoring and alerting, optimize infrastructure, and ensure reliable operations using custom-built operators and cdk8s.
  • Terraform – Implement Infrastructure-as-Code (IaC) to enable rapid provisioning, seamless configuration changes, and efficient scaling.
  • Automation & Development – Build and enhance automation tools and frameworks in Golang and Python to streamline operations.

On-Call Rotation

We maintain a structured on-call rotation to ensure 24/7 coverage:

Location & Eligibility

This is a remote role open to candidates located in US ideally located on Eastern or Central Time Zone.

RESPONSIBILITIES:  

  • Lead and mentor a small team of global Site Reliability Engineers
  • Ensure Reliability and Availability: You will ensure uptime for crucial services and systems based on business required SLOs. Minimize service disruptions through proactive monitoring, capacity planning and fault-tolerant design.
  • Architecture and System Design: you will design and architect complex, scalable and reliable systems.
  • Automation and Efficiency: you will develop and implement automation tools and frameworks to automate routine tasks to reduce human error and to streamline and improve operational processes to increase efficiency.
  • Build Observability and Monitoring tools: you will define, build, deploy, maintain, and extend our observability and monitoring tools to enhance system reliability and availability.
  • Incident Management and Response: you will maintain an effective on-call rotation to ensure 24/7 coverage. You will respond to incident response procedures to swiftly address and mitigate service disruptions.
  • Performance Monitoring and SLIs/SLOs: you will help define and monitor Service level Indicators (SLIs) and Service Level Objectives to set clear expectations for system performance.
  • Collaboration: you will work closely with product engineering to ensure service-level objectives and reliability targets are met
  • Problem-Solving & Troubleshooting: you respond to escalations by troubleshooting complex system and application incidents, perform root cause analysis, implement necessary corrective actions.
  • Thought Leadership and Innovation: Stay up to date with latest industry trends, emerging technologies. Iterate on best practices to increase the quality & velocity of development and deliverables.  

QUALIFICATIONS:   

  • 10+ years of experience in software engineering or site reliability engineering roles.
  • 2+ years of hands-on leadership role, managing and mentoring SRE engineering teams.
  • Proficiency in Golang and Python development skills
  • Extensive experience with cloud platforms (e.g., AWS, Azure, GCP) and cloud-native technologies.
  • Infrastructure-as-code (IaC): Deep understanding of Terraform core components (e.g., Terragrunt is a bonus) with real-world experience using Terraform for infrastructure provisioning and management.
  • Good knowledge of Kubernetes (e.g., cdk8s and operators are a bonus)
  • Solid experience developing Automation tools and frameworks.
  • Experience with Logging Solutions (e.g., Loki, Syslog, Elasticsearch, Logstash, Kibana, Filebeat, Fluentbit, etc.) 
  • Experience with Monitoring and Metrics Solutions (e.g., Prometheus, Grafana, Victoria Metrics)
  • Practical experience with Linux system administration
  • Experience with Version control system (e.g., Git, GitHub) and code review  
  •  Excellent communication skills are required.

US Pay Range

The US National annual base salary range for this full-time position is $244,000 – $265,000 + annual performance bonus

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

Aviatrix

View company profile →