Principal Data Scientist
OctusAbout the role
Octus
Octus is a leading global provider of credit intelligence, data, and analytics. Since 2013, tens of thousands of professionals across hedge fund, investment banking, management consulting, and law firm verticals have come to rely on Octus to make better, faster, and more confident decisions in pace with the fast-moving credit markets.
For more information, visit: https://octus.com/
Working at Octus
Octus hires growth-minded innovators and trailblazers across the globe to drive our business and culture. Our core values – Action Oriented, Customer First Mindset, Effective Team Players, and Driven to Excel – define an organizational ethos that’s as high-performing as it is human. Among other perks, Octus employees enjoy competitive health benefits, matched 401k and pension plans, PTO, generous parental leave, gym subsidies, educational reimbursements for career development, recognition programs, pet-friendly offices (US only), and much more.
Role
Job Description:
- Build pipelines for data acquisition by writing code for querying huge amounts of unstructured textual data from a variety of data sources like Financial SEC filings of publicly traded companies, private & public company press releases, co. transcripts, bond offering memorandum docs, etc. using Elasticsearch & RMySQL frameworks in R & database software like HeidiSQL to query databases like MySQL & MongoDB.
- Utilize frameworks like XML, rjson, pdftools in R to parse & process data from different sources & formats incl. .pdf, XML, json, csv etc. & store it in a structured & organized format for data processing, analysis, & modeling.
- Devise & implement processes to perform data preprocessing & assessment of data quality text processing & statistical techniques incl. imputation to handle missing data, data type conversions to maintain consistency in data integration, dimensionality reduction, normalization, feature aggregation, encoding, etc.
- Leverage frameworks in R incl. OpenNLP, Quanteda, tm, text2vec to provide comprehensive functionality for text analysis & natural language processing. Utilize frameworks daily for a variety of tasks incl. corpus creation & management, tokenization, formulation of doc. feature matrices, parts-of-speech tagging, entity extraction, etc. to generate analysis for data exploration, engineer features, formulate details of the model, & overall bld. robust frameworks for projects.
- Execute defined frameworks for project that req. data-driven solutions by building, execute, & test various data science models or enhance existing models using text mining & machine learning algorithms. Track & monitor model’s performance by testing & debugging when req. Incorporate feedback, business requests from stakeholders to continually improve & enhance workflow & performance.
- Conceptualize & build supervised &/or unsupervised models from structured &/or unstructured text data.
- Generate static & interactive data visualizations using frameworks & tools incl. ggplot, Shiny, d3.js to share & present complex ideas, results, project takeaways w/ tech. & non-tech. stakeholders.
- Review, evaluate, & communicate recommendations on modeling techniques & results to team, leadership, & stakeholders. Develop case studies using model output & suggest ways insights might be used.
- Deploy models in real-time by writing production-level code for scalable models & integrating it w/in the company’s data infrastructure.
- Collaborate & participate w/ different business units across the company to identify areas where data science can be used to automate manual processes.
- Mentor & lead new & junior members of the team.
- Formulate & implement ideas at intersection of distressed debt investing & data science, develop credit-risk models & transform into data products.
Education and Experience: Requires a Master’s degree in Data Science and 4 years of experience in job offered or 4 years of experience in the Related Occupation. Experience can be pre or post degree.
Related Occupation:
2 years of experience as a Data Scientist or any other job title performing the following job duties:
- Build pipelines for data acquisition by writing code for querying huge amounts of unstructured textual data from a variety of data sources like Financial SEC filings of publicly traded companies, private & public company press releases, co. transcripts, bond offering memorandum docs, etc. using Elasticsearch & RMySQL frameworks in R & database software like HeidiSQL to query databases like MySQL & MongoDB.
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s