Jobs and Careers
CA

Senior Site Reliability Engineer

Castleton Commodities International
United Statesfull_timeVerifiedPosted 6 Mar 2026

About the role

The Senior Site Reliability Engineer is responsible for improving the reliability, availability, scalability, and operational excellence of our critical infrastructure platforms and services. This role partners closely with Engineering, Security, and Infrastructure teams to design resilient cloud-native architectures, implement Infrastructure as Code (IaC) and CI/CD standards, and drive measurable reliability outcomes. The Senior Site Reliability Engineer will also lead efforts to define and validate recovery objectives (RTO/RPO), design and implement Business Continuity / Disaster Recovery (BCP/DR) plans, and coordinate structured testing to ensure readiness.

Responsibilities:

Reliability Engineering & Operations 

  • Own and improve service reliability through SLO/SLI definition, error budgets, and operational best practices. 

  • Design, implement, and maintain observability (monitoring, logging, tracing, alerting) to reduce MTTR and improve proactive detection. 

  • Lead incident response practices including on-call improvements, runbooks, post-incident reviews (RCA), and preventative actions. 

  • Partner with application teams to improve performance, capacity planning, and resiliency under failure scenarios. 

Infrastructure & Cloud Architecture 

  • Design and operate highly available, fault-tolerant Cloud architectures (multi-AZ and, where required, multi-region). 

  • Implement resilient patterns across compute, storage, networking, and managed services (e.g., autoscaling, load balancing, backups, replication). 

  • Drive cloud governance best practices (tagging, account/landing zone patterns, least privilege, guardrails) in partnership with security and platform teams. 

Infrastructure as Code (IaC) & DevOps Enablement 

  • Build and maintain IaC modules and standards (e.g., Terraform, CloudFormation, CDK) for repeatable, auditable infrastructure delivery. 

  • Develop, standardize, and optimize CI/CD pipelines to enable safe, automated deployments (e.g., GitHub Actions, GitLab CI, Jenkins, AWS CodePipeline). 

  • Promote DevOps practices: version-controlled infrastructure, automated testing, immutable deployments, and progressive delivery patterns. 

  • Establish environment consistency across dev/test/stage/prod and ensure infrastructure drift detection and remediation. 

BCP/DR, RTO/RPO Definition & Testing 

  • Collaborate with stakeholders to evaluate and define service-level RTO and RPO targets based on business and technical requirements. 

  • Design and implement BCP/DR architectures and procedures (backups, restore workflows, replication, failover/failback, data integrity validation). 

  • Coordinate and execute structured DR tests (tabletop, simulation, partial failover, full failover) and document outcomes. 

  • Maintain DR runbooks, dependency maps, and recovery checklists; drive remediation of gaps identified during testing. 

  • Produce metrics and reporting on DR readiness, test results, and continuous improvement actions. 

Qualifications:

  • 7+ years of experience in SRE, DevOps, Platform Engineering, or Systems Engineering roles supporting production environments. 

  • Strong proficiency with observability platforms (e.g., Datadog, Prometheus/Grafana, ELK/OpenSearch, Nagios, Nimsoft, etc). 

  • Strong hands

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

Castleton Commodities International

View company profile →