Jobs and Careers
NE

Senior Site Reliability Engineer

Nextiva
Spain (Remote), SpainRemotefull_timeVerifiedPosted 8 Oct 2025

About the role

<h3> </h3> <p><strong><span><span>Redefine the future of customer experiences. One conversation at a time.</span><br/></span></strong></p> <p>At Nextiva, we’re reimagining how businesses connect, bringing together customer experience and team collaboration on a single, conversation centric platform. Powered by AI, driven by human innovation.</p> <p>Our culture is forward thinking, customer obsessed and built on the belief that meaningful connections drive better business outcomes. Whether it’s through our signature Amazing Service®, the technology we create, or the experiences we cultivate, connection is at the core of who we are.</p> <p>If you’re ready to collaborate with incredible people, make an impact, and help businesses everywhere deliver truly amazing experiences, this is where you belong.</p> <p><span><strong>Build Amazing. Deliver Amazing. Live Amazing. Be Amazing.</strong></span></p>   <h3> </h3> <p> </p><p>We are looking for a Senior Site Reliability Engineer (SRE) to join our Middleware Engineering team. In this highly dynamic environment, you'll be responsible for supporting and scaling our Kafka and Elasticsearch infrastructure - core systems that power our SaaS platform.</p> <p>We're looking for someone who thrives on automation, embraces AI-driven observability, and is eager to learn and adopt new technologies quickly. You'll not only respond to production issues, but proactively build intelligent, resilient systems to prevent them.</p> <p>If you enjoy owning systems end to end, writing clean automation, and working in a fast-moving team that values innovation, this role is for you.</p> <p><strong>Key Responsibilities</strong> </p> <ul> <li>Triage, troubleshoot, and resolve complex production issues involving Kafka and Elasticsearch</li> <li>Design and build automated monitoring, alerting, and logging systems - leveraging AI/ML techniques where possible</li> <li>Write tools and infrastructure software to support self-healing, auto-scaling, and incident prevention</li> <li>Automate system administration tasks - from patching and upgrades to config and deployment workflows</li> <li>Use and manage GitHub extensively for infrastructure-as-code, release management, and collaboration</li> <li>Partner with development, QA, and performance teams to ensure middleware systems are production-ready</li> <li>Participate in the on-call rotation and continuously improve incident response and resolution playbooks</li> <li>Mentor junior engineers and contribute to a culture of automation, learning, and accountability</li> <li>Lead large-scale reliability and observability projects in collaboration with global teams</li> </ul> <p><strong>Qualifications </strong></p> <ul> <li>Bachelor's degree in Computer Science, Engineering, or equivalent practical experience</li> <li><span>Fluent English communication skills (spoken and written)</span></li> </ul> <p><strong>Core Competencies </strong></p> <ul> <li>6+ years of experience in software development, automation, or infrastructure engineering</li> <li>Deep experience with MongoDB, Kafka and/or Elasticsearch in production environments</li> <li>Strong Linux systems expertise and 6+ years managing Linux-based environments</li> <li>Hands-on experience with cloud platforms - GCP and/or AWS required</li> <li>Proficient in scripting languages like Python, Bash, etc</li> <li>Automation-first mindset - deep experience with Ansible, Terraform, Jenkins</li> <li>Expert-level understanding of Git and GitHub workflows for CI/CD and infrastructure-as-code</li> <li>Proficient with container tools (Docker) and orchestrators (Kubernetes)</li> <li>Strong understanding of SRE principles - SLAs/SLOs, alerting, observability, and incident management</li> <li>Experience with SQL, caching systems (e.g., Redis), and troubleshooting distributed systems</li> <li>Quick learner with a strong curiosity for new tools, frameworks, and AI/ML use cases in operations</li> </ul> <p><strong>Nice to Have </strong></p> <ul> <li>Observability Tools: Datadog, Splunk, Kibana, Opsgenie</li> <li>Programming: Java/Spring, JavaScript/React</li> <li>Middleware: RabbitMQ, Tomcat</li> <li>Experience with AI/ML-based anomaly detection, AIOps platforms, and LLM integrations for infrastructure</li> <li>Azure cloud experience (nice to have)</li> </ul> <p><strong> Why Join Us Why Join Us</strong></p> <ul> <li>Shape the future of middleware reliability using AI and intelligent automation</li> <li>Work with a global team that values initiative, innovation, and ownership</li> <li>Grow in a fast-paced environment where learning and experimentation are part of the culture</li> <li>Drive technical leadership, mentor others, and make a meaningful platform-wide impact</li> </ul> <p><strong>How to Apply</strong></p> <p>If you're passionate about automation, AIOps, MLOps, and scalable middleware infrastructure, and you're ready to move fast, learn constantly, and own critical systems - we'd love to connect with you.</p> <p><s

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

Nextiva

View company profile →