Manager, DevOps Engineer
SCORAbout the role
This senior individual contributor role is accountable for architecture-level decisions, implementation, and operational ownership of secure, resilient AWS cloud-native platform capabilities supporting Velogica US within SCOR Digital Solutions. The position leads modernization of legacy scripts, tooling, and release patterns into Terraform/CloudFormation-based Infrastructure as Code, modern CI/CD, and standardized operational guardrails across development, QA, staging, Rescore, and production environments. The role partners closely with engineering, architecture, and security stakeholders to improve delivery speed, reliability, and compliance, and carries deep ownership of Amazon RDS/Aurora lifecycle management, performance, recovery readiness, and database-aware deployment coordination for business-critical services.
Cloud-Native Platform Engineering (AWS)
- Design and operate resilient AWS platform capabilities across dev, QA, staging, Rescore, and production.
- Build secure, reusable cloud patterns for networking, compute, storage, and identity aligned to enterprise standards.
- Lead migration of legacy operational workflows into cloud-native, event-driven, and API-first automation.
- Partner with architecture and security teams to align platform implementation with governance and risk controls.
- Serve as a senior technical owner for platform standards, reference architectures, and complex cross-team delivery decisions.
Infrastructure as Code & Environment Provisioning
- Own Infrastructure as Code strategy and implementation using Terraform (preferred) and/or AWS CloudFormation.
- Build reusable modules/templates for VPC, IAM, compute, container platforms, observability, and data services.
- Enforce immutable, repeatable environment provisioning through Git-based workflows and automated validation.
- Eliminate manual runbooks and shell-heavy operations by codifying infrastructure and operational procedures.
Modern CI/CD & Release Engineering
- Design and maintain modern CI/CD pipelines using GitHub Actions and/or GitLab CI (Jenkins de-emphasized/legacy support only).
- Implement pipeline standards for build, test, security scanning, artifact promotion, and deployment orchestration.
- Enable progressive delivery approaches (rolling, canary, blue-green) with automated rollback and health validation.
- Improve deployment lead time, reliability, and change safety with standardized release controls and telemetry.
AI-Assisted Engineering & Developer Productivity
- Use AI coding assistants (especially GitHub Copilot in VS Code) to accelerate modernization of legacy scripts, pipelines, and infrastructure code.
- Collaborate on practical guardrails for AI-assisted code generation, review, testing, and security validation in engineering workflows.
- Partner with teams to adopt prompt patterns and coding standards that improve delivery speed while preserving code quality and compliance.
- Identify and scale high-value AI-assisted refactoring and documentation use cases across platform and application repositories.
Containerization & Runtime Operations
- Build and operate containerized deployment patterns using Docker and AWS container platforms (ECS/EKS).
- Define runtime standards for configuration, secrets, service discovery, scaling, and resilience.
- Partner with application teams to modernize service packaging and runtime architecture for cloud-native workloads.
RDS, Database Reliability, & Data Platform Operations
- Maintain deep ownership of Amazon RDS/Aurora operations, lifecycle management, and production reliability.
- Plan and execute database upgrades, patching, scaling, backups, restore testing, and performance tuning.
- Lead database-sensitive deployment coordination, schema change planning, and environment refresh/rebuild activities.
- Support data migration and replication patterns (including AWS DMS where appropriate).
- Define and automate guardrails for database observability, capacity, and recovery objectives (RTO/RPO).
- Provide senior oversight for database risk, change sequencing, and production readiness across release events.
Observability, Incident Response, & SRE Practices
- Implement comprehensive monitoring, alerting, logging, and tracing using CloudWatch and approved enterprise tooling.
- Define SLIs/SLOs and operational dashboards for platform and service health.
- Participate in on-call and incident response; lead root cause analysis and preventive remediation.
- Improve operational excellence through post-incident learning, automation, and reliability engineering
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s