Senior Cloud Engineer
AHEADAbout the role
The Senior Cloud Engineer - Managed Services is responsible for leading the day-to-day operation, administration, monitoring, support, and continuous improvement of customer cloud environments, including production Red Hat OpenShift (Kubernetes) platforms, across public Clouds such as Azure, AWS, GCP or OCI. This role is primarily focused on advanced cloud operations, ITSM execution, customer advisory, mentorship, and operational improvement, with secondary exposure to automation, DevOps, and reliability practices where they improve consistency, scale, and service quality. The role is expected to independently own complex and high-risk operational work, serve as a senior escalation point for major incidents and changes, and partner with architects on major design or platform engineering decisions as appropriate. There is an expectation of some travel and after-hours/weekend support in the event of major outages/issues, or when requested by the client for change windows (or similar events).
Duties & Responsibilities:
- Lead the support and operation of cloud infrastructure, platform services, identity, networking, security controls, and operational tooling across customer environments.
- Able to architect and lead deployment of moderately complex solutions related to cloud solutions.
- Understands performance, scaling and functional characteristics of software technologies
- Ability to understand open-source and cloud use-cases and recommend standard design patterns commonly used in such solutions (best practices).
- Own complex incidents, escalations, and problem investigations; perform advanced troubleshooting, coordination, service restoration, and follow-through to durable resolution.
- Plan and execute complex changes and recurring operational activities including provisioning, access changes, maintenance events, backup and recovery validation, patching coordination, and platform hygiene.
- Serve as a senior escalation point within the on-call rotation for major incidents, high-impact issues, and customer-approved after-hours change activity.
- Follow and reinforce established ITSM processes for incident, request, change, problem, escalation, documentation, and customer-facing status communication.
- Develop and maintain runbooks, SOPs, standards, knowledge articles, and technical documentation that improve consistency and service quality.
- Mentor other Cloud Engineers, review work for quality and completeness, and provide technical guidance on operational best practices.
- Drive monitoring, alerting, logging, tagging, policy, compliance, and cost-visibility improvements that strengthen managed cloud operations.
- Use scripting, automation, and AI, to reduce repetitive effort, improve consistency, and scale service delivery.
- General familiarity with DevOps/SRE tooling is required but is not the primary emphasis of the role.
- Participate in customer meetings, service reviews, and advisory discussions; translate technical issues, risk, and improvement opportunities into clear business-facing communication.
- Operate and support Red Hat OpenShift (Kubernetes) clusters in production, including cluster health, upgrades, scaling, and lifecycle management.
- Manage OpenShift access and security controls, including RBAC, SCCs, NetworkPolicies, secrets management, and certificate/ingress considerations.
- Troubleshoot platform and workload issues across Kubernetes/OpenShift constructs (nodes, operators, routes/ingress, services, deployments, pods, persistent volumes) and coordinate remediation with application, network, and security teams.
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s