Jobs and Careers
ST

AVP, CL, Production Eng, WRB Tech

Standard Chartered
Bukit Jalil KL, MYfull_timeVerifiedPosted 19 Aug 2026

About the role

Key Responsibilities

Service stability and incident management
•    Ensure maximum service quality and stability through prompt and effective response to technical incidents.
•    Act as a catalyst for change by performing incident and problem analysis, identifying root causes, and driving continual service improvement (CSI) initiatives.
•    Where relevant, perform a control function to ensure that new technology changes do not introduce instability into the production environment.
Monitoring and observability
•    Own and drive the achievement of “north star” monitoring and observability goals.
•    Ensure comprehensive monitoring, alerting, and logging are in place for critical services, enabling proactive detection and rapid remediation of issues.
Automation and operational excellence
•    Lead the automation of operational tasks such as deployments, monitoring, scaling, and infrastructure management to reduce manual effort and operational risk.
Site Reliability Engineering (SRE) practices
•    Participate in and oversee incident response, troubleshooting, and post-incident reviews (post-mortems) to minimise downtime and institutionalise learning from failures.
•    Optimise infrastructure, systems, and processes for performance, efficiency, and reliability.
•    Contribute to the design and implementation of robust deployment pipelines and release strategies that enable smooth, frequent, and reliable releases (e.g. blue/green, canary).
Change, release, and rollout management
•    Review production-related changes, releases, and rollouts with zero or minimal impact to application stability and client experience.
•    Review and coordinate dependent changes across surrounding systems, infrastructure, networks, and shared services.
•    Ensure thorough technical plans are in place for all production changes, including implementation steps, fallback/rollback strategies, data conversion or migration plans, and validation checks.

Reporting and continuous improvement
•    Provide inputs for monthly dashboards and reports, including incident and problem trends, key service metrics, and the status of Service Improvement Plans (SIPs) and Root Cause Analysis (RCA) action items.
•    Track and drive closure of remediation actions to prevent recurrence of incidents.
Collaboration, coaching, and knowledge sharing
•    Participate in and support cross-training and structured knowledge transfer activities within and across support and engineering teams.
•    Promote SRE and production engineering best practices across the chapter and wider organisation, fostering a culture of shared ownership for reliability and operational excellence.
Leverage AI and automation for production engineering
•    Use AI-driven tools (e.g. for log analysis, anomaly detection, alert correlation, and capacity forecasting) to proactively identify, diagnose, and resolve production issues.
•    Collaborate with engineering and platform teams to integrate AI/ML capabilities into monitoring, incident management, and self-healing workflows (e.g. automated remediation, intelligent runbooks).
•    Continuously review and refine AI-enabled alerts, models, and automations based on production behaviour, incident learnings, and feedback from support teams.
•    Promote the safe and compliant adoption of AI solutions within production engineering, ensuring adherence to the bank’s risk, security, and data governance standards.
Strategy
•    To be accountable to execute the strategy devised for the business unit 
Business
•    Fully accountable in incident, problem, change, and risk management for production applications/systems, including:

Incident Management
•    Owns end-to-end management of all production incidents impacting the application/system, from detection and triage through to resolution and closure.
•    Coordinates technical and business stakeholders during incidents to ensure timely communication, clear ownership, and rapid restoration of service.
•    Ensures incidents are correctly classified, prioritised, and logged, with accurate d

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

Standard Chartered

View company profile →