Jobs and Careers
IN

Site Reliability Engineer-I

Innovaccer
Indiafull_timeVerifiedPosted 12 Apr 2023

About the role

<h4><em><strong>Your Role</strong></em></h4> <p>As an <strong><em>Site Reliability Engineer-I</em></strong>, you will be responsible for building/automating secure cloud Infrastructure (Infrastructure As A Code - IaaC) with various pillars Cost, Reliability, Scalability, Performance, Cost, Deployment, Service Availability - SLA/SLO/SLI, Performance etc.</p> <p><strong>A Day in the Life</strong></p> <ul> <li>Build CICD stack collaborating across Dev and QA/Automation team and drive organization to new level of (daily/hourly) continuous delivery and deployment.</li> <li>Security is paramount to everything we do, you will work closely with CISO, Dev team(s) and make security as first class citizens. Develop S-CICD (Secure CICD), enable various security tool chains and vulnerability reports to developers via automation.</li> <li>Observability is very critical for the scale of our systems and ability to find insights/behavior, detect problem/failures. Looking for leads to drive this charter spanning across logs, metrics, mesh, tracing etc.</li> <li>Collaborate closely with Dev and QA team to bring given initiative to a closer, increase adoption of DevOps practices and tool chain.</li> <li>Apply strong analytical skills to understand production system metrics, drive change, optimize system utilization and drive cost efficiency.</li> <li>Autoscale/down the platform during peak season scenarios.</li> <li>Understand end to end platform architecture and how to best and fast perform triage/RCA by looking at various data points derived from observability tool chain.</li> <li>You will be part of the <strong>24x7 OnCall Production</strong> Support team.</li> <li>Lead monthly operations review with the executive team. Some examples include, but are not limited to – Platform/Application/Infrastructure KPIs - UpTime, RCA , CAP <br/>(Corrective Action Plan) and PAP (Preventive Action Plan), security reports, audit reports.</li> <li>You will be responsible for Operating and Managing production and staging cloud platforms, responsible for Ops (executing/automation runbook/SOP/ Maintain <br/>up-time/SLA) as well as Site Reliability engineering.</li> <li>Ensure that the Platform is secured as per guidelines established by CISO. e,g, Secure against DDoS attacks by implementing WAF, Vulnerability and Patch management, install required security agents etc.</li> <li>Lead least privilege based RBAC for various production services and tool chains.</li> <li>Build and execute Disaster Recovery plan.</li> <li>Key stakeholder to participate in case of IR (Incident Response).</li> </ul> <p> </p> <p><strong>What You Need</strong></p> <ul> <li>Proven work experience of 1-4 years in DevOps/SRE.</li> <li>Solid experience with at least one of the clouds with automation focus - <strong>AWS, Azure, GCP</strong>. Certification has advantages.</li> <li>Hands-on experience with Kubernetes along with Linux.</li> <li>Programming experience with scripting languages e.g. Python.</li> <li>Build and deployment experience building scalable CICD architectures and solutions is preferred.</li> <li>Building observability stack from logs, metrics, traces, service mesh, data observability is preferred.</li> <li>Building reliability, scalability and performance systems in Production. This requires significant engineering experience and risk evaluation.</li> <li>Good at documenting and structuring documents for consumption by various dev teams.</li> <li>Experience working in a Production environment with process focus is preferred.</li> <li>Ticketing system, Incident management experience is preferred.</li> <li>Cloud Security is a major advantage and highly preferred skill.</li> <li>Hands-on experience with a few of these - Kafka, Postgres, SnowFlake etc. is preferred.</li> <li>Bachelor’s Degree or equivalent.</li> </ul> <p><strong>Personality Trait</strong></p> <ul> <li>Able to perform with cool head under pressure situations without taking any shortcuts.</li> <li>Collaboration with solid verbal and oral communication skills are very critical to this role. Possesses excellent verbal and written communication skills and the ability to interact professionally with a diverse group of developers, product owners, and subject matter experts.</li> <li>Strong cross-functional collaboration skills, relationship building skills, and ability to achieve results without direct reporting relationships</li> <li>Ability to quickly identify and drive to the optimal solution when presented with a series of constraints.</li> <li>Excellent judgment, analytical thinking, and problem-solving skills.</li> <li>Self-motivated individual that possesses excellent time management and organizational skills.</li> <li>Strong sense of personal responsibility and accountability for delivering high quality work.</li> </ul> <p><strong>Preferred Skills:</strong></p> <ul> <li>MultiCloud - AWS, Azure, GCP</li> <li>Distributed Compute - Kubernetes (EKS/AKS), Containerization</li> <li>Persistence store

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

Innovaccer

View company profile →