Software Development Engineer- Product Reliability Engineering
VisaAbout the role
Company Description
Visa is a world leader in payments and technology, with over 259 billion payments transactions flowing safely between consumers, merchants, financial institutions, and government entities in more than 200 countries and territories each year. Our mission is to connect the world through the most innovative, convenient, reliable, and secure payments network, enabling individuals, businesses, and economies to thrive while driven by a common purpose – to uplift everyone, everywhere by being the best way to pay and be paid.
Make an impact with a purpose-driven industry leader. Join us today and experience Life at Visa.
Visa’s Technology Organization is a community of problem solvers and innovators reshaping the future of commerce. We operate the world’s most sophisticated processing networks capable of handling more than 65k secure transactions a second across 80M merchants, 15k Financial Institutions, and billions of everyday people. While working with us you’ll get to work on complex distributed systems and solve massive scale problems centered on new payment flows, business and data solutions, cyber security, and B2C platforms.
Job Description
Every time someone taps, swipes, or clicks to pay- Visa infrastructure makes it happen in milliseconds, across 200+ countries. As a Software Development Engineer on the Product Reliability Engineering (PRE) team, you won’t just watch those systems run- you’ll be one of the engineers building, automating, and evolving them.
PRE is not a traditional ops team. We are a software engineering organization that treats infrastructure as code, reliability as a product, and automation as a strategic advantage. You’ll write Python, build agentic AI tools, manage data platforms, and contribute to the distributed systems that process billions of real-time transactions. From day one, you are an engineer- and from day one, your work matters.
If you are endlessly curious about how large-scale systems stay resilient, obsess over elegant automation, and want to launch your career at the intersection of AI, infrastructure, and global financial technology — this role was built for you.
Build Automation That Scales
▪ Design and ship end-to-end automation for deployment pipelines, infrastructure provisioning, and release orchestration — code that runs millions of times so engineers never have to repeat themselves.
▪ Write clean, production-grade Python (and Go or Bash where it counts) to eliminate toil, reduce manual intervention, and make systems self-managing.
▪ Develop modular frameworks for release scheduling, validation, rollback, and reporting that integrate across the full software delivery lifecycle.
Manage & Evolve Data Platforms
▪ Support the build, deployment, and operations of relational database systems, contributing to schema design, architecture decisions, and solution engineering for critical payment data infrastructure.
▪ Gain exposure to real-time event streaming architectures that support payment processing at scale
▪ Perform database health operations including patching, upgrades, backups, and recovery to maintain the availability and integrity of tier-1 production databases.
▪ Optimize query performance through index tuning, execution plan analysis, and replication monitoring — targeting metrics like query execution time, CPU usage, and replication latency.
▪ Automate database tasks and configuration management using tools like Ansible and Liquibase, and contribute to CI/CD pipelines that govern schema changes through TEST and PROD environments safely.
▪ Build predictive and reactive monitoring dashboards for database anomalies, surfacing health signals before they become incidents.
Ship Agentic AI & ML-Powered Tools
▪ Build GenAI-powered engineering assistants that automate deployment orchestration, release governance, and environment lifecycle management.
▪ Integrate LLMs into observability, incident response, and developer support workflows, transforming reactive operations into proactive, AI-driven intelligence.
▪ Contribute to prompt engineering, model fine-tuning, and agentic automation initiatives that position PRE as one of the most AI-forward reliability organizations in financial technology.
Own Observability & Platform Health
▪ Build dashboards, alerts, and metrics using Prometheus, Grafana, Splunk, or ELK that give engineers real-time clarity on complex, globally distributed systems.
▪ Analyze system performance and availability data and turn insights into infrastructure improvements that prevent incidents before they occur.
▪ Contribute to self-healing and auto-scaling capabilities that keep critical payment infrastructure resilient without human interventi
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s