Staff Site Reliability Engineer
CookUnityAbout the role
About CookUnity
Food has lost its soul to modern convenience. And with it, has lost the power to nourish, inspire, and connect us. So in 2018, CookUnity was founded as the first-of-its-kind platform that connects the world with the source of truly great food: chefs. Today, CookUnity delivers 35 million meals a year from the industry’s best chefs to homes all over the country. Fresh. Ready-to-eat. And crafted with the passion that nourishes body and soul.
Unwilling to stop there, CookUnity is expanding beyond delivery to become an ever-innovating marketplace focused on our singular mission: empower Chefs to nourish the world.
If that mission has you hungry in more ways than one, you’ve found the right job posting.
About the Team:
The CookUnity Infrastructure team is responsible for maintaining our highly available infrastructure that services our millions of customers, guaranteeing availability, reliability, and confidentiality. The team services the requests of the engineering organization related to CICD pipelines, builds, infrastructure, and security.
The role:
The DevOps Engineer is responsible for architecting, implementing, and maintaining robust cloud-native infrastructure and deployment pipelines with a focus on reliability, scalability, and automation. This role requires experience with AWS, Kubernetes (EKS), ArgoCD, GitHub, and familiarity with development languages such as Kotlin, Python, and Bash. The engineer will collaborate closely with software development and operations teams to ensure continuous delivery, system reliability, and rapid incident response in a dynamic environment.
This opportunity is open in: Brazil, Chile, Colombia, Mexico, Peru and Argentina.
Responsibilities:
- Architect, deploy, and manage highly available and scalable infrastructure on AWS, leveraging services such as EC2, VPC, S3, IAM, and EKS.
- Design, implement, and maintain Kubernetes clusters (EKS) and oversee the deployment of containerized applications using best practices for security, scaling, and automation.
- Develop and manage GitOps workflows using ArgoCD for automated, reliable, and auditable application deployments to Kubernetes.
- Write and maintain infrastructure as code (IaC) using tools such as Terraform.
- Build, optimize, and troubleshoot CI/CD pipelines to support rapid, reliable software delivery, integrating with ArgoCD and other modern DevOps toolchains.
- Develop robust automation scripts and tools in languages such as Kotlin, Python, and/or Bash to streamline operational processes, monitoring, and incident response.
- Proactively monitor system performance, reliability, and security, responding to incidents and participating in on-call rotations as needed.
- Collaborate with software engineers to improve deployment strategies, system observability, and overall site reliability.
- Implement and enforce security best practices across all infrastructure and deployment workflows.
- Maintain comprehensive documentation of infrastructure, processes, and procedures for operational transparency and team knowledge sharing.
- Experience using GitHub and GitHub Actions to automate, testing and deployments.
Qualifications:
- 7+ years in DevOps, SRE, or related roles in cloud-native environments, with at least 5 years of direct experience managing AWS infrastructure at scale
- Proficiency in deploying, managing, and troubleshooting Kubernetes clusters, especially AWS EKS, including networking, RBAC, and Helm.
- Advanced English Level
- Advanced hands-on experience with ArgoCD for GitOps-based Kubernetes deployments, including setup, configuration, and troubleshooting.
- Strong development and scripting skills in Kotlin, Python, and Bash, with the ability to build automation tools and integrate with APIs.
- Deep knowledge of CI/CD concepts and tools, with proven experience building and maintaining pipelines for cloud-native applications.
- Demonstrated ability to design and implement infrastructure as code using Terraform and/or AWS CloudFormation.
- Strong problem-solving skills, including root cause analysis and incident management in distributed, cloud-based systems,
- Excellent communication and collaboration abilities, working effectively across development, QA, and operations teams.
Preferred requirements:
- Bachelor’s or Master’s degree in Computer Science, Engineering, or related field or equivalent experience.
- Relevant certifications (e.g., AWS Certified DevOps Engineer, Certified Kubernetes Administrator).
- Experience with additional DevOps tools such as Jenk
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s