Senior Reliability Engineer/ Senior Systems Administrator
AmericorAbout the role
Americor is seeking a Senior Reliability Engineer/Senior Systems Administrator to join our team. A Senior Site Reliability Engineer is a highly skilled technical professional responsible for ensuring the reliability, scalability, and performance of complex cloud-based systems and applications by actively monitoring infrastructure, automating operational tasks, identifying potential issues, troubleshooting incidents, and collaborating with development teams to implement solutions that prioritize system.
About Americor
Americor is a leader in debt relief solutions, helping thousands of clients achieve financial freedom through innovative services. We are looking for a Site Reliability Engineer (SRE) to support and enhance the infrastructure of our mission-critical CRM project, ensuring seamless service delivery to help people regain control of their financial futures. As a recognized ‘Top Place to Work’ and ‘Best Company’ in customer service, quality, and value, we value collaboration, growth, and innovation in every team member.
Project Features
- Hosted in OVHcloud (US)
- OVHcloud contains Bare Metal and VMs
- OS: CentOS / AlmaLinux OS
- Components: Nginx, KeyDB/Redis, OpenSearch, RabbitMQ
- Database: MariaDB/MySQL, Percona / Galera Cluster, ProxySQL, Maxscale
- Storage: GlusterFS/NFS/Ceph
- Networking: HAProxy, VyOS, iptables
- Language: PHP 8 (PHP-FPM, Yii2, Symfony, Laravel, OPcache)
- Monitoring tools: Datadog, Vector, Sentry
- IaC: Ansible, Terraform
- Alerting: OpsGenie/Jira Service Management
Responsibilities:
- Ensure the reliability of infrastructure supporting mission-critical services, minimizing downtime and optimizing performance.
- Proactively monitor, respond to, diagnose, and resolve incidents, improving response time and minimizing customer impact.
- Work closely with Russian-speaking developers, as well as QA and system analysts.
- Enhance CI/CD pipelines, monitoring tools, and automation processes to streamline workflows and increase system efficiency.
- Keep infrastructure-related documentation up to date.
Requirements:
- 5+ years of experience in a Site Reliability Engineering role, with a proven track record of maintaining high-availability infrastructure in a high-load environment.
- Expertise in Linux systems and web stacks (Nginx, PHP, MySQL/MariaDB, Redis/KeyDB) to ensure smooth and efficient operation.
- Strong experience with MySQL/MariaDB Galera cluster and Gluster storage to optimize data reliability and scalability.
- Deep knowledge of network architectures, including TCP/IP, DNS, VPNs, and load-balancing techniques, with hands-on experience in troubleshooting and optimizing network performance to support distributed systems across multiple regions.
- Proficiency in PHP and Docker for seamless integration and deployment of services.
- Solid understanding of CI/CD and security best practices to drive efficiency and protect our systems.
- Experience establishing a GitOps workflow using Infrastructure-as-Code (Terraform) and Configuration-as-Code (Ansible) to provision infrastructure and configurations for high-load environments.
- Understanding the principles of Infrastructure-as-Code, Monitoring-as-Code, and GitOps (we use Ansible and Terraform).
- Experience with Cloudflare and AWS services (EKS, S3, OpenSearch).
- Experience building fault-tolerant systems and compliance audits (SOC, FFIEC, etc.).
- Familiarity with Jira and Agile software development.
- Familiarity with modern container orchestration and deployment tools (Kubernetes, Helm).
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s