Production Services Specialist ll
Bank of AmericaAbout the role
Job Description:
At Bank of America, we are guided by a common purpose to help make financial lives better through the power of every connection. We do this by driving Responsible Growth and delivering for our clients, teammates, communities and shareholders every day.
Being a Great Place to Work is core to how we drive Responsible Growth. This includes our commitment to being an inclusive workplace, attracting and developing exceptional talent, supporting our teammates’ physical, emotional, and financial wellness, recognizing and rewarding performance, and how we make an impact in the communities we serve.
Bank of America is committed to an in-office culture with specific requirements for office-based attendance and which allows for an appropriate level of flexibility for our teammates and businesses based on role-specific considerations.
At Bank of America, you can build a successful career with opportunities to learn, grow, and make an impact. Join us!
Job Description:
This job is responsible for providing front-line support to end users, responding to issues related to incidents and problem management governance for multiple applications, and leading triage activities on all business impacting incidents. Key responsibilities include ensuring compliance with incident management and problem management policies and procedures, serving as a focal point for the customer, client, and associate experience, restoring complex production incidents under tight Service Level Agreements, and pursuing root cause and problem resolution follow ups.
Position Summary
Hadoop Production Platform Operations & Support: This Production support specialist role for managing multi-tenant shared Enterprise Hadoop production platform operations which has close to 200 applications and 14 different clusters. This Role requires niche skill set of Hadoop Administration to monitor and manage the cluster. This role is responsible for the cluster availability to perform BAU operations by application tenants and user analytics. The resource has to coordinate effectively with all the support groups, Vendor and user base to ensure prompt evaluation of a situation, and efficient resolution in a production environment. This role is also responsible for the resiliency and ensure to conduct the Disaster recovery exercises on timely manner facilitating the resources required by the shared tenants. Identify the areas of automation for the team and keep track of the implementation.
Responsibilities:
- Leads production support triage efforts, manages bridge line troubleshooting, engages in technical research, and escalates issues to leadership as needed
- Ensures all impacts are accurately recorded and documented in the system of record, oversees that documents and wikis are updated and available for use during triage, and supports the documentation of application flows, upstream/downstream impacts during outages, the customer experience, and contacts for support needs
- Identifies and/or validates business impacts through interpretation of monitors, dashboards, and logs to communicate with leadership and vendors
- Manages activities to identify incident root cause, resolution, preventative actions, and change requests, and reports on incident data quality
- Promotes and enforces production governance during triage/testing and identifies production failure scenarios, vulnerabilities, and opportunities for improvement
- Serves as a subject matter expert for applications within a portfolio, leveraging extensive knowledge of application functionalities and application flows
- Assesses and prioritizes research requests, ad hoc reports, and offline incidents at the direction of senior team members and delegates work as needed to team members and peers
Required Qualifications
- 5+ years in technology experience.
- Hadoop System Administration skills with at least 3+ years of experience of handling Cloudera distribution of Hadoop or Horton works.
- Experience on evaluating the Proof of Concept on new Hadoop tools and related technologies. Formulate and design systems scope and objectives for applications and the development of information technology projects.
- Extensive experience in big data tools: Hadoop, Hive, Impala and Spark.
- Proficiency in Scala, Python, SQL, and PySpark. Strong experience and triaging skills on the hive, Spark data analysis & Impala.
- Experience with stream-processing systems: Kafka, Spark-Streaming, etc.
- Experience on configuring high availability of the Cloudera Hadoop services to achieve never down situation on Hadoop components.
- Good understanding of Operating System (OS) concepts, process management, capacity
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s