Principal Data Scientist - Customer Data Machine Learning Team
Capital OneAbout the role
Data is at the center of everything we do. As a startup, we disrupted the credit card industry by individually personalizing every credit card offer using statistical modeling and the relational database, cutting edge technology in 1988! Fast-forward a few years, and this little innovation and our passion for data has skyrocketed us to a Fortune 200 company and a leader in the world of data-driven decision-making.
As a Data Scientist at Capital One, you’ll be part of a team that’s leading the next wave of disruption at a whole new scale, using the latest in computing and machine learning technologies and operating across billions of customer records to unlock the big opportunities that help everyday people save money, time and agony in their financial lives.
Team Description
The Enterprise Consumer Data Science team develops customer matching models that are foundational to servicing our customers and internal data consumers across all Capital One lines of business. The team develops and deploys entity resolution systems operating on massive datasets (hundreds of millions of records) in both batch and realtime, leveraging distributed compute and algorithm optimization to overcome scaling challenges. In this team, you will contribute to the full lifecycle of model development, from scoping business requirements to deployment and monitoring on Capital One's kubernetes platform, and your work will enable a variety of use cases including customer servicing, anti-money laundering, marketing, and bankruptcy processing. Our models are required to make impactful decisions in the face of uncertainty, and we leverage a variety of state-of-the-art interpretable and trustworthy machine learning techniques to support transparency. The team is also heavily invested in expanding core ML technology stack in deep-learning/GenAI and graph analytics and embedding research.
In this role, you will
Partner with a cross-functional team of data scientists, software engineers, and product managers to deliver a product customers love
Leverage a broad stack of technologies — Python, Conda, AWS, H2O, Spark, and more — to reveal the insights hidden within huge volumes of numeric and textual data
Build machine learning models through all phases of development, from design through training, evaluation, validation, and implementation
Flex your interpersonal skills to translate the complexity of your work into tangible business goals
The Ideal Candidate is
Innovative. You continually research and evaluate emerging technologies. You stay current on published state-of-the-art methods, technologies, and applications and seek out opportunities to apply them.
Creative. You thrive on bringing definition to big, undefined problems. You love asking questions and pushing hard to find answers. You’re not afraid to share a new idea.
Technical. You’re comfortable with open-source languages and are passionate about developing further. You have hands-on experience developing data science solutions using open-source tools and cloud computing platforms.
Statistically-minded. You’ve built models, validated them, and back tested them. You know how to interpret a confusion matrix or a ROC curve. You have experience with clustering, classification, sentiment analysis, time series, and deep learning.
A data guru. "Big data" doesn't faze you. You have the skills to retrieve, combine, and analyze data from a variety of sources and structures. You know understanding the data is often the key to great data science.
Basic Qualifications
Currently has, or is in the process of obtaining a Bachelor’s Degree plus 5 years of experience in data analytics, or currently has, or is in the process of obtaining a Master’s Degree plus 3 years in data analytics, or currently has, or is in the process of obtaining PhD, with an expectation that required degree will be obtained on or before the scheduled start date
At least 1 year of experience in open source programming languages for large scale data analysis
At least 1 year of experience with machine learning
At least 1 year of experience with relational databases
Preferred Qualifications
Master’s Degree in “STEM” field (Science, Technology, Engineering, or Mathematics)
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s