EverCommerce - Senior Site Reliability Engineer
EverCommerceAbout the role
EverCommerce (Nasdaq: EVCM) is a leading service commerce platform, providing vertically-tailored, integrated SaaS solutions that help more than 690,000 global service-based businesses accelerate growth, streamline operations, and increase retention. Its modern digital and mobile applications create predictable, informed, and convenient experiences between customers and their service professionals. With its EverPro, EverHealth, and EverWell brands specializing in Home, Health, and Wellness service industries, EverCommerce provides end-to-end business management software, embedded payment acceptance, marketing technology, and customer experience applications. Learn more at EverCommerce.com.
We are building an extraordinary company and looking for talented, energetic, and motivated people to join our team. You can learn more about our Company, Culture and Values here: https://www.evercommerce.com/about-us/careers/
We are looking for a Systems Engineer to focus on our Cloud Engineering Operations team.
The Senior Site Reliability Engineer role is responsible for ensuring the reliability, performance, and operational stability of business-critical SaaS infrastructure across cloud and traditional hosting environments. This position independently supports production infrastructure, collaborates closely with Senior Cloud Engineers and cross-functional technology teams, and leads implementation of infrastructure improvements, operational automation, monitoring enhancements, and modernization initiatives. The ideal candidate is an experienced infrastructure engineer with strong technical skills across Windows, Linux, virtualization, networking, automation, and AWS who enjoys solving complex operational challenges while continuously improving the reliability of production systems.
Responsibilities:
- Maintain uptime, reliability, and performance for production SaaS environments across AWS, colocation, and hosted infrastructure platforms.
- Independently support Windows and Linux infrastructure, virtualization platforms, storage systems, and networking components.
- Lead implementation of infrastructure modernization and operational automation initiatives while partnering with Senior Cloud Engineers on larger strategic efforts.
- Troubleshoot complex infrastructure, application connectivity, and production incidents while leading root cause analysis and recommending long-term corrective actions.
- Design and implement improvements to monitoring, alerting, and operational visibility across infrastructure platforms.
- Collaborate with security teams and technology leaders to support SOX, PCI, and HIPAA compliance initiatives.
- Lead vulnerability remediation efforts and infrastructure lifecycle maintenance activities.
- Develop and maintain technical documentation, operational procedures, and infrastructure standards.
- Collaborate with application teams and vendors to support reliable infrastructure hosting production database systems.
- Independently lead medium-sized infrastructure projects from planning through implementation.
- Participate in disaster recovery testing, recovery planning, and continuous operational improvement initiatives.
- Evaluate emerging technologies and recommend practical improvements to infrastructure operations.
Skills and Experience needed for success in this role:
Required:
- 8 years of systems, infrastructure, or cloud engineering experience supporting production environments.
- Experience administering Windows Server and Linux systems.
- Experience supporting infrastructure in AWS environments.
- Experience with enterprise virtualization platforms.
- Experience designing or implementing infrastructure automation using PowerShell, Bash, Python, or similar scripting languages.
- Experience leading technical investigations and root cause analysis for production issues.
- Understanding of networking fundamentals, firewalls, backup and recovery processes, and operational resiliency.
- Familiarity with enterprise monitoring platforms and operational troubleshooting.
- Familiarity with operational support for production database infrastructure, including backup validation, connectivity troubleshooting, and recovery operations.
- Strong communication, collaboration, and technical documentation skills.
- Ability to independently manage complex production infrastructure with minimal supervision.
Preferred:
- Experience with Terraform or Infrastructure as Code.
- Experience with configuration management technologies such as Ansible, Puppet, or Chef.
- Experience s
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s