Machine Learning Architect
U.S. BankAbout the role
At U.S. Bank, we’re on a journey to do our best. Helping the customers and businesses we serve to make better and smarter financial decisions and enabling the communities we support to grow and succeed. We believe it takes all of us to bring our shared ambition to life, and each person is unique in their potential. A career with U.S. Bank gives you a wide, ever-growing range of opportunities to discover what makes you thrive at every stage of your career. Try new things, learn new skills and discover what you excel at—all from Day One.
Job Description
At U.S. Bank, we are committed to leveraging industry-leading technology to enhance our financial services. Our goal is to empower customers and businesses with advanced tools for smarter financial decision-making and to support community growth through innovative solutions. A career with U.S. Bank offers a dynamic environment where you can explore innovative technologies, develop technical skills, and excel in your field from Day One.
SUMMARY
We are seeking a highly skilled and experienced Machine Learning Architect to lead the development and optimization of advanced Data Science and AI Machine Learning systems. This role demands in-depth expertise in systems, GPU, performance engineering, enterprise security, big data, secure and scalable ML infrastructure, and advanced AI technologies. You will play a pivotal role in architecting and implementing reliable AI infrastructure to ensure scalability, security, and high-performance solutions.
JOB DESCRIPTION
The Machine Learning Architect is responsible for designing, optimizing, securing, and documenting enterprise-scale AI and data systems for the end-to-end technical integrity, performance, and reliability. This includes system-level performance tuning, robust ML infrastructure, and cloud architecture. The role ensures that all AI workflows meet enterprise standards for scalability, security, and operational excellence, while enabling innovation through AI technologies and best practices. This role is critical in bridging deep technical expertise with enterprise-scale AI execution, ensuring that AI systems are not only powerful but also reliable, secure, and future-ready.
KEY RESPONSIBILITIES
System Design and Implementation
-Design and implement scalable, high-performance architectures for machine learning, data science, and AI workflows.
- Develop programs to automate workflows, deployments, and monitoring for improved operational efficiency.
- Optimize GPU utilization and performance using CUDA and algorithmic optimization to accelerate machine learning workloads.
System Maintenance and Troubleshooting
- Diagnose and resolve issues related to Linux servers, networks, GPUs, cluster health, job failures, and performance bottlenecks to ensure system reliability and efficiency.
- Drive performance improvements across ML pipelines, inference engines, model serving frameworks (e.g., Triton, vLLM), and data processing layers.
- Develop and maintain reliable AI workflows with robust failover mechanisms to ensure uninterrupted operations in production environments.
Security and Compliance
- Implement enterprise-grade IAM authentication and authorization with Azure SSO, OIDC, Kerberos, and Active Directory, for ML APIs, model endpoints, and data platforms.
Migration and Integration
- Migrate existing AI applications on Cloudera Hadoop/Hive/Spark to Azure cloud services, ensuring data integrity, performance tuning, and validation.
Documentation and Collaboration
- Create and maintain documentation for system configurations, operational procedures, and troubleshooting knowledge bases to support knowledge sharing.
- Collaborate with cross-functional teams to translate business requirements into technical solutions and provide technical leadership and mentorship.
- Stay updated on emerging technologies in AI, machine learning, Big Data, system clustering, and enterprise security, and recommend their adoption where appropriate.
BASIC QUALIFICATIONS
- Advanced degree in Computer Science, Engineering, or related field.
- Deep expertise in AI technologies, synthetic data, automation, advanced analytics.
- 10+ years of hands-on experience in systems engineering, ML infrastructure, and performance optimization.
- Proficiency in Linux, clustering, and distributed systems.
- Expertise in GPU monitoring, GPU scheduling, CUDA, algorithmic optimization, and parallel computing.
- Proficiency in languages such as shell, Ansible, C/C++, Golang, Java, and Python for automating workflows, deployments, and monitoring.
- Deep understanding of Deep Learning, Computer Vision, LLMs, vector databases, and AI
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s