Jobs and Careers
DE

Junior Site Reliability Engineer

Deel
Anywhere (LATAM)Remotefull_timeVerifiedPosted 4 Jul 2025

About the role

Who we are is what we do.

Deel is the all-in-one payroll and HR platform for global teams. Our vision is to unlock global opportunity for every person, team, and business. Built for the way the world works today, Deel combines HRIS, payroll, compliance, benefits, performance, and equipment management into one seamless platform. With AI-powered tools and a fully owned payroll infrastructure, Deel supports every worker type in 100+ countries—helping businesses scale smarter, faster, and more compliantly.

Among the largest globally distributed companies in the world, our team of 5,000 spans more than 100 countries, speaks 74 languages, and brings a connected and dynamic culture that drives continuous learning and innovation for our customers.

Why should you be part of our success story?

As the fastest-growing Software as a Service (SaaS) company in history, Deel is transforming how global talent connects with world-class companies – breaking down borders that have traditionally limited both hiring and career opportunities. We're not just building software; we're creating the infrastructure for the future of work, enabling a more diverse and inclusive global economy. In 2024 alone, we paid $11.2 billion to workers in nearly 100 currencies and provided healthcare and benefits to workers in 109 countries—ensuring people get paid and protected, no matter where they are.

Our momentum is reflected in our achievements and customer satisfaction: CNBC Disruptor 50,  Forbes Cloud 100, Deloitte Fast 500, and repeated recognition on Y Combinator’s top companies list – all while maintaining a 4.83 average rating from 15,000 reviews across G2, Trustpilot, Captera, Apple and Google.

Your experience at Deel will be a career accelerator. At the forefront of the global work revolution, you'll tackle complex challenges that impact millions of people's working lives. With our momentum—backed by a $12 billion valuation and $800  million in Annual Recurring Revenue (ARR) in just over five years—you'll drive meaningful impact while building expertise that makes you a sought-after leader in the transformation of global work.

Summary
The Site Reliability Engineer (SRE) plays a critical role in ensuring the high reliability, scalability, and performance of Deel's systems, with a strong focus on automation, efficiency, and continuous improvement. Combining software engineering practices with operations, you will manage, scale, and optimize production services, prioritizing resilience and user experience. You will work closely with Development, Infra, and DevOps teams to proactively monitor, analyze, and resolve issues, directly impacting Deel's ability to innovate and maintain high-quality services for its global customers.

Responsibilities

  • Monitor production systems for performance and reliability issues, and respond to incidents swiftly (using tools like Datadog, Prometheus, Grafana, Loki, Zabbix).

  • Establish, track, and maintain Service Level Objectives (SLOs), Service Level Agreements (SLAs), and Service Level Indicators (SLIs) for critical services.

  • Build and maintain automation scripts, tools, and dashboards to improve monitoring, alerting, and response times.

  • Handle incidents including:

    • Triage & Escalation: Quickly assess incidents, mitigate impact, and escalate when necessary.

    • Root Cause Analysis: Conduct post-mortems and root cause analysis, implementing corrective measures to prevent recurrence with the teams.

    • Incident Documentation: Maintain thorough documentation for each incident, including timelines, impact, and resolution steps.

    • Analytics: Provide weekly, monthly and yearly analytics for incidents, problems, Mean Time To Repair, and uptime metrics.

  • Implement robust alerting and escalation protocols for production incidents, with escalation paths to minimize downtime. Validate alerts to ensure accurate thresholds before production use.

  • Analyze data from production systems for bottlenecks, performance issues, and optimization opportunities, using insights from dev/staging to identify trends and preemptively address potential issues.

  • Identify inefficiencies and bottlenecks in system processes and work to resolve them.

  • Assist in n building run books and recommend automated recovery procedures.

  • Integrate lessons learned from incidents and team retrospectives into workflows and processes.

  • Regularly collaborate with product teams to understand user needs, ensuring that improvements align with customer experience goals.

Qualifications

  • 1+ years of relevant Site Reliability Engineering or production operations experience

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

Deel

View company profile →