Jobs and Careers
VI

Staff Site Reliability Engineer - Hadoop Administration

Visa
United Statesfull_timeVerifiedPosted 19 Apr 2023

About the role

Company Description

Visa is a world leader in digital payments, facilitating more than 215 billion payments transactions between consumers, merchants, financial institutions and government entities across more than 200 countries and territories each year. Our mission is to connect the world through the most innovative, convenient, reliable and secure payments network, enabling individuals, businesses and economies to thrive.

When you join Visa, you join a culture of purpose and belonging – where your growth is priority, your identity is embraced, and the work you do matters. We believe that economies that include everyone everywhere, uplift everyone everywhere. Your work will have a direct impact on billions of people around the world – helping unlock financial access to enable the future of money movement.

Join Visa: A Network Working for Everyone.

Job Description

Product Reliability Engineering (PRE) is part of the Visa's O& I Technology organization. The division is responsible for maintaining and supporting Visa's data assets and provides support for value added products and services to drive innovation for our partners and clients, within Visa and globally. Product Reliability Engineering Big Data Platform Team is part of PRE supports open-source Big Data and Kafka clusters in Visa.

As a Staff Big Data Engineer, you will be responsible for monitoring, troubleshooting, automating and continuously developing software product and tools to improve the availability and resiliency of open-source Big Data Platforms at Visa.  In this hands-on role, you will Administer and ensure performance, reliability and increase the operational efficiency of open-source big data platforms.

Essential Function

  • Effective, polished interaction with customer to capture information quickly, explain customer responsibilities in resolving issue, communicate next steps and status, and inspire confidence
  • Person will be responsible to Perform Big Data Administration and Engineering activities on multiple opensource Hadoop, Kafka, HBase and Spark clusters
  • Strong Troubleshooting and debugging skills
  • Cross-team teamwork, build and maintain relationships with the customer teams, the user community, architects, and engineering teams, jointly work on key deliverables ensuring production scalability and stability
  • Effective Root cause analysis of major production incidents and developing learning documentation
  • Identify and implement HA solution for services with SPOF.
  • Plan and perform capacity expansion and upgrades in timely manner avoiding any scaling issues and bugs.
  • Automation of repetitive tasks to reduce manual effort and avoid Human errors.
  • Tune alerting and setup observability to proactively identify the issues and performance problems.
  • Work closely with L-3 teams in reviewing new use cases, cluster hardening techniques for building a robust and reliable platform.
  • Create SOP documents and guidelines on effectively managing and utilizing the platforms.
  • Leverage Devops tools, disciplines(Incident, problem and change management) and standards in day to operations.
  • Ensure the Hadoop platform can effectively meet performance and SLA requirements.
  • Perform automation, selfheal, develop software tools and products as per the requirement.

This is a hybrid position. Hybrid employees can alternate time between both remote and office. Employees in hybrid roles are expected to work from the office 2-3 set days a week (determined by leadership/site), with a general guidepost of being in the office 50% or more of the time based on business needs.

Qualifications

Basic Qualifications

5+ years of relevant work experience and a Bachelors degree, OR 8+ years of relevant work experience


Preferred Qualification

6 or more years of work experience with a bachelor’s degree or 4 or more years of relevant experience with an advanced degree.
Hands on experience working as a Hadoop system engineer in managing Hadoop platforms.
Experience in building, managing and tuning performance of Hadoop platforms.
Extensive knowledge on Hadoop eco-system such as Zookeeper, HDFS, Yarn, HIVE, SPARK and Kafka.
Excellent Shell, Python programming skills for automation requirement for repetitive dev-ops tasks
Understanding of security tools like Kerberos and Ranger.
Experience on Hortonworks distribution or Open Source preferred
Hands-on experience in debugging Hadoop issues both on platform and applications.
Knowledge on HBASE and Kubernetes is a plus.
Good understanding of Linux, networking, CPU, memory and storage.
Excellent interpersonal, verbal, and written communication skills.

Additional Information

Work Hours: Varies upon the needs of the departme

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

Visa

View company profile →