Senior Site Reliability Engineer
Castleton Commodities InternationalAbout the role
Responsibilities:
Reliability Engineering & Operations
Own and improve service reliability through SLO/SLI definition, error budgets, and operational best practices.
Design, implement, and maintain observability (monitoring, logging, tracing, alerting) to reduce MTTR and improve proactive detection.
Lead incident response practices including on-call improvements, runbooks, post-incident reviews (RCA), and preventative actions.
Partner with application teams to improve performance, capacity planning, and resiliency under failure scenarios.
Infrastructure & Cloud Architecture
Design and operate highly available, fault-tolerant Cloud architectures (multi-AZ and, where required, multi-region).
Implement resilient patterns across compute, storage, networking, and managed services (e.g., autoscaling, load balancing, backups, replication).
Drive cloud governance best practices (tagging, account/landing zone patterns, least privilege, guardrails) in partnership with security and platform teams.
Infrastructure as Code (IaC) & DevOps Enablement
Build and maintain IaC modules and standards (e.g., Terraform, CloudFormation, CDK) for repeatable, auditable infrastructure delivery.
Develop, standardize, and optimize CI/CD pipelines to enable safe, automated deployments (e.g., GitHub Actions, GitLab CI, Jenkins, AWS CodePipeline).
Promote DevOps practices: version-controlled infrastructure, automated testing, immutable deployments, and progressive delivery patterns.
Establish environment consistency across dev/test/stage/prod and ensure infrastructure drift detection and remediation.
BCP/DR, RTO/RPO Definition & Testing
Collaborate with stakeholders to evaluate and define service-level RTO and RPO targets based on business and technical requirements.
Design and implement BCP/DR architectures and procedures (backups, restore workflows, replication, failover/failback, data integrity validation).
Coordinate and execute structured DR tests (tabletop, simulation, partial failover, full failover) and document outcomes.
Maintain DR runbooks, dependency maps, and recovery checklists; drive remediation of gaps identified during testing.
Produce metrics and reporting on DR readiness, test results, and continuous improvement actions.
Qualifications:
7+ years of experience in SRE, DevOps, Platform Engineering, or Systems Engineering roles supporting production environments.
Strong proficiency with observability platforms (e.g., Datadog, Prometheus/Grafana, ELK/OpenSearch, Nagios, Nimsoft, etc).
Strong hands
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s