Vice President, AI / Machine Learning Software Engineer
BNYAbout the role
At BNY, our culture allows us to run our company better and enables employees’ growth and success. As a leading global financial services company at the heart of the global financial system, we influence nearly 20% of the world’s investible assets. Every day, our teams harness cutting-edge AI and breakthrough technologies to collaborate with clients, driving transformative solutions that redefine industries and uplift communities worldwide.
Recognized as a top destination for innovators, BNY is where bold ideas meet advanced technology and exceptional talent. Together, we power the future of finance – and this is what #LifeAtBNY is all about. Join us and be part of something extraordinary.
Role Overview
We are seeking a senior‑level engineer to design, build, and operate production‑grade GenAI and Retrieval‑Augmented Generation (RAG) platforms at scale. This role focuses on industrializing LLM‑based systems with strong guardrails, observability, evaluation frameworks, and operational rigor, ensuring reliability, safety, and cost efficiency across the full AI lifecycle. This role is located in Jersey City, NJ.
Key Responsibilities
- Design and build production‑ready RAG pipelines, including retrieval, ranking, prompt orchestration, and response generation, with comprehensive guardrails, tracing, and observability.
- Implement offline and online evaluation frameworks for prompts, models, and datasets, including quality, safety, latency, and cost metrics.
- Own end‑to‑end lifecycle management for GenAI systems, covering prompt versions, model versions, datasets, and configurations.
- Establish and maintain CI/CD pipelines for prompts, models, and data, enabling safe, repeatable, and auditable releases.
- Implement cost and performance monitoring, including token usage, inference latency, throughput, and spend optimization.
- Build and enforce safety mechanisms, such as content filtering, policy enforcement, red‑teaming feedback loops, and abuse detection.
- Define and operationalize incident management workflows, including alerting, triage, rollback mechanisms, and post‑incident analysis.
- Partner closely with product, platform, and governance teams to ensure GenAI solutions meet enterprise reliability, security, and compliance standards.
- Mentor engineers and influence best practices for building scalable, trustworthy AI systems.
What Success Looks Like
- GenAI systems that are observable, measurable, and resilient, not “black boxes.”
- Safe and cost‑efficient RAG pipelines running reliably in production.
- Fast iteration cycles with strong controls, enabling teams to ship GenAI features with confidence.
Required Qualifications
- Advanced degree in STEM engineering degree, or equivalent work experience with experience preferred in related fields. 7-9 years of related experience required; experience in the securities or financial services industry is a plus
- Strong experience building and operating production ML or GenAI systems in enterprise environments.
- Deep hands‑on expertise with LLM orchestration frameworks, such as LangChain and/or LlamaIndex.
- Experience with model registries and experiment tracking, such as MLflow or equivalent.
- Solid understanding of Kubernetes‑based deployments and cloud‑native architectures.
- Familiarity with feature stores, data pipelines, and retriever/index lifecycle management.
- Proven experience implementing telemetry, logging, metrics, and distributed tracing for ML/AI workloads.
- Strong knowledge of CI/CD practices for ML, GenAI, and data‑driven systems.
Preferred Qualifications
- Experience operating LLM systems at scale, including multi‑model or multi‑provider strategies.
- Exposure to AI safety, governance, and compliance frameworks in regula
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s