Jobs and Careers
SE

Staff AWS Site Reliability Engineer

SearchStax
US - Remote, United StatesRemotefull_timeVerifiedPosted 5 Sept 2025
💰 $240,000/yr($170,000/yr$240,000/yr)

About the role

About us

SearchStax is a leading cloud-native search platform enabling web teams to deliver powerful search in an easy, fast, and cost-effective way. We are on a mission to make powerful search easy for enterprises across the globe. We are self-funded and profitable.

Our products are used by 600+ brand-name customers. The search market is growing fast. We feel we are uniquely positioned to continue to lead the search market for many years to come.

Our team is composed of smart, driven subject matter experts who love to collaborate and solve problems in new / creative ways. We value the importance of bringing diverse backgrounds and interests to the collaboration process. We prioritize work-life balance and strive to promote an energizing and healthy environment.

Our Values

  • Ownership

  • Lead humbly

  • Results focused

  • Customer Obsession

  • Embrace and drive change

  • Innovation and continual Improvement

About the Role

We are looking for an experienced Staff AWS Site Reliability Engineer (SRE) to join our Infrastructure team. This is a hands-on leadership role for someone who thrives in fast-paced startup environments and has experience scaling infrastructure to support large, high-traffic, mission-critical systems. You will be responsible for driving reliability, scalability, and performance across our platform, ensuring that our infrastructure grows with our customers and business.

If this sounds like you, let’s talk!

What You Will Do

  • Lead and Own Scalability & Reliability: Take ownership of scaling our AWS infrastructure to support thousands of servers and high-growth workloads.

  • Automation First: Design and implement automation frameworks for provisioning, monitoring / logging, scaling, and recovery, minimizing manual operations.

  • Performance Optimization: Continuously evaluate and tune systems for latency, throughput, and cost efficiency.

  • Reliability Engineering: Build systems that are resilient, self-healing, and observable—using SLOs, error budgets, and reliability best practices.

  • Cross-Functional Collaboration: Partner closely with Development, QA, and Product Engineering teams to deliver highly available and performant systems.

  • Incident Management: Own on-call processes, lead root-cause analysis, and implement preventive measures.

  • Mentorship & Leadership: Act as a technical leader in the team, mentoring other engineers and setting standards for best practices.

Why Join Us

SearchStax is at an inflection point of growth. We are scaling fast, with ambitious plans to expand our infrastructure capacity by 15x over the next 2–3 years. As part of this journey, you’ll:

  • Play a critical role in building a world-class, scalable, and reliable infrastructure that powers search for enterprises and institutions around the globe.

  • Be part of a startup-minded team where your expertise and decisions directly shape the company’s trajectory.

  • Work on hard problems at scale—from automation to observability to multi-region architectures.

  • Join a company that is trusted by 1,200+ global customers across higher education, healthcare, and enterprise, with deep partnerships in the DXP ecosystem.

This is an opportunity to leave your mark—leading infrastructure that must scale exponentially and building systems that will support SearchStax’s growth for years to come.

What You Must Have

  • Experience: 7+ years in Site Reliability, DevOps, or Infrastructure Engineering roles, with startup experience and a track record of scaling infrastructure to thousands of servers.

  • Deep AWS Expertise: Hands-on mastery of AWS services (EC2, EKS, RDS, S3, CloudFront, VPC, IAM, etc.), with strong understanding of multi-region, highly available architectures.

  • Infrastructure as Code: Proficiency in Terraform, CloudFormation, or similar tools.

  • Automation & Scripting: Strong skills in Python, Go, or similar languages for building automation.

  • Monitoring & Observability: Expertise with tools such as Prometheus, Grafana, Loki, ELK/EFK, Datadog, or equivalent.

  • CI/CD &

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

SearchStax

View company profile →