Jobs and Careers
NE
Senior Site Reliability Engineer
NextivaSpain (Remote), SpainRemotefull_timeVerifiedPosted 8 Oct 2025
About the role
<h3> </h3>
<p><strong><span><span>Redefine the future of customer experiences. One conversation at a time.</span><br/></span></strong></p>
<p>At Nextiva, we’re reimagining how businesses connect, bringing together customer experience and team collaboration on a single, conversation centric platform. Powered by AI, driven by human innovation.</p>
<p>Our culture is forward thinking, customer obsessed and built on the belief that meaningful connections drive better business outcomes. Whether it’s through our signature Amazing Service®, the technology we create, or the experiences we cultivate, connection is at the core of who we are.</p>
<p>If you’re ready to collaborate with incredible people, make an impact, and help businesses everywhere deliver truly amazing experiences, this is where you belong.</p>
<p><span><strong>Build Amazing. Deliver Amazing. Live Amazing. Be Amazing.</strong></span></p>
<h3> </h3>
<p> </p><p>We are looking for a Senior Site Reliability Engineer (SRE) to join our Middleware Engineering team. In this highly dynamic environment, you'll be responsible for supporting and scaling our Kafka and Elasticsearch infrastructure - core systems that power our SaaS platform.</p>
<p>We're looking for someone who thrives on automation, embraces AI-driven observability, and is eager to learn and adopt new technologies quickly. You'll not only respond to production issues, but proactively build intelligent, resilient systems to prevent them.</p>
<p>If you enjoy owning systems end to end, writing clean automation, and working in a fast-moving team that values innovation, this role is for you.</p>
<p><strong>Key Responsibilities</strong> </p>
<ul>
<li>Triage, troubleshoot, and resolve complex production issues involving Kafka and Elasticsearch</li>
<li>Design and build automated monitoring, alerting, and logging systems - leveraging AI/ML techniques where possible</li>
<li>Write tools and infrastructure software to support self-healing, auto-scaling, and incident prevention</li>
<li>Automate system administration tasks - from patching and upgrades to config and deployment workflows</li>
<li>Use and manage GitHub extensively for infrastructure-as-code, release management, and collaboration</li>
<li>Partner with development, QA, and performance teams to ensure middleware systems are production-ready</li>
<li>Participate in the on-call rotation and continuously improve incident response and resolution playbooks</li>
<li>Mentor junior engineers and contribute to a culture of automation, learning, and accountability</li>
<li>Lead large-scale reliability and observability projects in collaboration with global teams</li>
</ul>
<p><strong>Qualifications </strong></p>
<ul>
<li>Bachelor's degree in Computer Science, Engineering, or equivalent practical experience</li>
<li><span>Fluent English communication skills (spoken and written)</span></li>
</ul>
<p><strong>Core Competencies </strong></p>
<ul>
<li>6+ years of experience in software development, automation, or infrastructure engineering</li>
<li>Deep experience with MongoDB, Kafka and/or Elasticsearch in production environments</li>
<li>Strong Linux systems expertise and 6+ years managing Linux-based environments</li>
<li>Hands-on experience with cloud platforms - GCP and/or AWS required</li>
<li>Proficient in scripting languages like Python, Bash, etc</li>
<li>Automation-first mindset - deep experience with Ansible, Terraform, Jenkins</li>
<li>Expert-level understanding of Git and GitHub workflows for CI/CD and infrastructure-as-code</li>
<li>Proficient with container tools (Docker) and orchestrators (Kubernetes)</li>
<li>Strong understanding of SRE principles - SLAs/SLOs, alerting, observability, and incident management</li>
<li>Experience with SQL, caching systems (e.g., Redis), and troubleshooting distributed systems</li>
<li>Quick learner with a strong curiosity for new tools, frameworks, and AI/ML use cases in operations</li>
</ul>
<p><strong>Nice to Have </strong></p>
<ul>
<li>Observability Tools: Datadog, Splunk, Kibana, Opsgenie</li>
<li>Programming: Java/Spring, JavaScript/React</li>
<li>Middleware: RabbitMQ, Tomcat</li>
<li>Experience with AI/ML-based anomaly detection, AIOps platforms, and LLM integrations for infrastructure</li>
<li>Azure cloud experience (nice to have)</li>
</ul>
<p><strong> Why Join Us Why Join Us</strong></p>
<ul>
<li>Shape the future of middleware reliability using AI and intelligent automation</li>
<li>Work with a global team that values initiative, innovation, and ownership</li>
<li>Grow in a fast-paced environment where learning and experimentation are part of the culture</li>
<li>Drive technical leadership, mentor others, and make a meaningful platform-wide impact</li>
</ul>
<p><strong>How to Apply</strong></p>
<p>If you're passionate about automation, AIOps, MLOps, and scalable middleware infrastructure, and you're ready to move fast, learn constantly, and own critical systems - we'd love to connect with you.</p>
<p><s
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s