Jobs and Careers
IN
Site Reliability Engineer-I
InnovaccerIndiafull_timeVerifiedPosted 12 Apr 2023
About the role
<h4><em><strong>Your Role</strong></em></h4>
<p>As an <strong><em>Site Reliability Engineer-I</em></strong>, you will be responsible for building/automating secure cloud Infrastructure (Infrastructure As A Code - IaaC) with various pillars Cost, Reliability, Scalability, Performance, Cost, Deployment, Service Availability - SLA/SLO/SLI, Performance etc.</p>
<p><strong>A Day in the Life</strong></p>
<ul>
<li>Build CICD stack collaborating across Dev and QA/Automation team and drive organization to new level of (daily/hourly) continuous delivery and deployment.</li>
<li>Security is paramount to everything we do, you will work closely with CISO, Dev team(s) and make security as first class citizens. Develop S-CICD (Secure CICD), enable various security tool chains and vulnerability reports to developers via automation.</li>
<li>Observability is very critical for the scale of our systems and ability to find insights/behavior, detect problem/failures. Looking for leads to drive this charter spanning across logs, metrics, mesh, tracing etc.</li>
<li>Collaborate closely with Dev and QA team to bring given initiative to a closer, increase adoption of DevOps practices and tool chain.</li>
<li>Apply strong analytical skills to understand production system metrics, drive change, optimize system utilization and drive cost efficiency.</li>
<li>Autoscale/down the platform during peak season scenarios.</li>
<li>Understand end to end platform architecture and how to best and fast perform triage/RCA by looking at various data points derived from observability tool chain.</li>
<li>You will be part of the <strong>24x7 OnCall Production</strong> Support team.</li>
<li>Lead monthly operations review with the executive team. Some examples include, but are not limited to – Platform/Application/Infrastructure KPIs - UpTime, RCA , CAP <br/>(Corrective Action Plan) and PAP (Preventive Action Plan), security reports, audit reports.</li>
<li>You will be responsible for Operating and Managing production and staging cloud platforms, responsible for Ops (executing/automation runbook/SOP/ Maintain <br/>up-time/SLA) as well as Site Reliability engineering.</li>
<li>Ensure that the Platform is secured as per guidelines established by CISO. e,g, Secure against DDoS attacks by implementing WAF, Vulnerability and Patch management, install required security agents etc.</li>
<li>Lead least privilege based RBAC for various production services and tool chains.</li>
<li>Build and execute Disaster Recovery plan.</li>
<li>Key stakeholder to participate in case of IR (Incident Response).</li>
</ul>
<p> </p>
<p><strong>What You Need</strong></p>
<ul>
<li>Proven work experience of 1-4 years in DevOps/SRE.</li>
<li>Solid experience with at least one of the clouds with automation focus - <strong>AWS, Azure, GCP</strong>. Certification has advantages.</li>
<li>Hands-on experience with Kubernetes along with Linux.</li>
<li>Programming experience with scripting languages e.g. Python.</li>
<li>Build and deployment experience building scalable CICD architectures and solutions is preferred.</li>
<li>Building observability stack from logs, metrics, traces, service mesh, data observability is preferred.</li>
<li>Building reliability, scalability and performance systems in Production. This requires significant engineering experience and risk evaluation.</li>
<li>Good at documenting and structuring documents for consumption by various dev teams.</li>
<li>Experience working in a Production environment with process focus is preferred.</li>
<li>Ticketing system, Incident management experience is preferred.</li>
<li>Cloud Security is a major advantage and highly preferred skill.</li>
<li>Hands-on experience with a few of these - Kafka, Postgres, SnowFlake etc. is preferred.</li>
<li>Bachelor’s Degree or equivalent.</li>
</ul>
<p><strong>Personality Trait</strong></p>
<ul>
<li>Able to perform with cool head under pressure situations without taking any shortcuts.</li>
<li>Collaboration with solid verbal and oral communication skills are very critical to this role. Possesses excellent verbal and written communication skills and the ability to interact professionally with a diverse group of developers, product owners, and subject matter experts.</li>
<li>Strong cross-functional collaboration skills, relationship building skills, and ability to achieve results without direct reporting relationships</li>
<li>Ability to quickly identify and drive to the optimal solution when presented with a series of constraints.</li>
<li>Excellent judgment, analytical thinking, and problem-solving skills.</li>
<li>Self-motivated individual that possesses excellent time management and organizational skills.</li>
<li>Strong sense of personal responsibility and accountability for delivering high quality work.</li>
</ul>
<p><strong>Preferred Skills:</strong></p>
<ul>
<li>MultiCloud - AWS, Azure, GCP</li>
<li>Distributed Compute - Kubernetes (EKS/AKS), Containerization</li>
<li>Persistence store
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s