Jobs and Careers
AV

Staff Engineer - Site Reliability

Aviatrix
United States, United Statesfull_timeVerifiedPosted 3 Feb 2025
💰 $190,000/yr($177,000/yr$190,000/yr)

About the role

The Aviatrix SRE team is a small but highly skilled global group of Systems Engineers/SREs dedicated to ensuring the reliability, availability, and performance of Aviatrix’s critical systems and services. Our mission is to build and maintain a robust, resilient infrastructure that enables Aviatrix to deliver high-quality services with agility through automation, best practices, and a culture of operational excellence.

About the Role

As an SRE – Staff Engineer, you’ll play a key role in designing, implementing, and maintaining highly available, fault-tolerant, and scalable systems. You’ll focus on automation, proactive monitoring, and Infrastructure-as-Code (IaC) to drive efficiency and reliability across our services.

Tech Stack & Responsibilities

  • Kubernetes – Manage application lifecycles, automate operational tasks, troubleshoot issues, integrate monitoring and alerting, optimize infrastructure, and ensure reliable operations using custom-built operators and cdk8s.
  • Terraform – Implement Infrastructure-as-Code (IaC) to enable rapid provisioning, seamless configuration changes, and efficient scaling.
  • Automation & Development – Build and enhance automation tools and frameworks in Golang and Python to streamline operations.

On-Call Rotation

We maintain a structured on-call rotation to ensure 24/7 coverage:

During Business Hours (rotates every 2 days)

  • EST: 9 AM – 6 PM
  • CST: 8 AM – 5 PM
  • PST: 6 AM – 3 PM

Outside Business Hours (6 PM – 9 AM PT, rotates weekly: Monday to Monday)

Location & Eligibility

This is a remote role open to candidates located in the US or Canada. You must be eligible to work in either country and currently reside there.

If you're passionate about building resilient infrastructure, automating operations, and ensuring system reliability at scale, we'd love to hear from you! 🚀

RESPONSIBILITIES:  

  • Ensure Reliability and Availability: You will ensure uptime for crucial services and systems based on business required SLOs. Minimize service disruptions through proactive monitoring, capacity planning and fault-tolerant design.
  • Architecture and System Design: you will design and architect complex, scalable and reliable systems.
  • Automation and Efficiency: you will develop and implement automation tools and frameworks to automate routine tasks to reduce human error and to streamline and improve operational processes to increase efficiency.
  • Build Observability and Monitoring tools: you will define, build, deploy, maintain, and extend our observability and monitoring tools to enhance system reliability and availability.
  • Incident Management and Response: you will maintain an effective on-call rotation to ensure 24/7 coverage. You will respond to incident response procedures to swiftly address and mitigate service disruptions.
  • Performance Monitoring and SLIs/SLOs: you will help define and monitor Service level Indicators (SLIs) and Service Level Objectives to set clear expectations for system performance.
  • Collaboration: you will work closely with product engineering to ensure service-level objectives and reliability targets are met
  • Problem-Solving & Troubleshooting: you respond to escalations by troubleshooting complex system and application incidents, perform root cause analysis, implement necessary corrective actions.
  • Thought Leadership and Innovation: Stay up to date with latest industry trends, emerging technologies. Iterate on best practices to increase the quality & velocity of development and deliverables.  

QUALIFICATIONS:   

  • 8+ years of experience maintaining and deploying highly available, fault-tolerant systems at scale. 
  • Proficiency in Golang or Python is required.
  • Infrastructure-as-code (IaC): Deep understanding of Terraform core components (e.g., Terragrunt is a bonus) with real-world experience using Terraform for infrastructure provisioning and management.
  • At least one cloud service provider experience (e.g., AWS, GCP, Azure, OCI)  
  • Good knowledge with Kubernetes (e.g., cdk8s and operators are a bonus)
  • Solid experience developing Automation tools and frameworks.
  • Experience with Logging Solutions (e.g., Loki, Syslog, Elasticsearch, Logstash, Kibana, Filebeat, Fluentbit, etc.) 
  • Experience with Monitoring and Metrics Solutions (e.g., Prometheus, Grafana, Victoria Metrics)
  • Practical experience with Linux system administration
  • Experience with Version con

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

Aviatrix

View company profile →