Lead Devops Engineer ( GCP & FedRAMP)
VirtusaAbout the role
Description
Job Description
The Lead DevOps Engineer, a key member of the EIT DevOps Team, is responsible for the staging and production infrastructure of Iron Mountain’s Digital Services within the federal sector. This role is pivotal in managing and optimizing staging and production deployment environments across Google Cloud Platform (GCP), Amazon Web Services (AWS), and Microsoft Azure.
Core responsibilities include provisioning and maintaining secure, scalable, and robust cloud infrastructure for the InSight DXP Platform. The Senior DevOps Engineer will apply extensive knowledge of cloud services and DevOps best practices to ensure application efficiency, high availability, and performance.
Additionally, this role involves creating and maintaining FedRAMP controls and documentation compliance. The Senior DevOps Engineer will execute automation pipelines, upgrade infrastructure, troubleshoot complex issues, and contribute to the ongoing enhancement of deployment processes. Close collaboration with development, operations, and other EIT teams is crucial for delivering seamless and reliable solutions.
Core Responsibilities:
- Cloud Infrastructure Management: Deploy, manage, and maintain cloud infrastructure across AWS, Azure, and/or GCP, ensuring compliance for government workloads.
- Infrastructure Automation: Automate infrastructure provisioning using Infrastructure as Code (IaC) tools like Terraform, OpenTofu, or AWS CloudFormation.
- Deployment Pipeline Streamlining: Collaborate with development teams to streamline CI/CD pipelines using tools such as GitLab and OpenTofu for efficient infrastructure and application delivery.
- Performance Optimization: Monitor system performance, participate in capacity planning, and optimize application and infrastructure performance by tuning configurations and identifying bottlenecks.
- Automation Development: Develop scripts and tools to automate routine operations, including patching, scaling, and monitoring.
- Self-Healing Systems: Design and implement self-healing systems that proactively detect and resolve faults.
- Data Integrity & Availability: Manage backup and disaster recovery strategies to ensure data integrity and availability across environments.
- Security & Compliance: Perform regular security audits and vulnerability patching, adhering to government compliance requirements (e.g., FedRAMP, NIST).
Incident Management & Observability:
- Real-time Incident Resolution: Respond to and resolve infrastructure incidents and outages in real-time, minimizing disruption.
- Root Cause Analysis (RCA): Conduct RCA for production issues and implement long-term corrective actions.
- On-Call Participation: Participate in an on-call rotation, escalating and coordinating responses to high-severity issues.
- Incident Documentation: Document incidents, responses, and postmortems to capture lessons learned.<
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s