Senior Software Engineer- Data Science Platform- Generative AI Inference
BloombergAbout the role
Description & Requirements
Bloomberg runs on data. It's our business and our product. From the biggest banks to elite hedge funds, financial institutions need timely, accurate data to capture opportunities and evaluate risk in fast-moving markets. With petabytes of data available, a solution to transform and analyze the data is critical to our success.
Bloomberg’s Data Science Platform was established to support development efforts around data-driven science, machine learning, and business analytics. The solution aims to provide scalable compute, specialized hardware and first-class support for a variety of workloads such as ML training job and inference services, Spark, and Jupyter. The solution was developed to provide a standard set of tooling for addressing the Model Development Life Cycle from experimentation and training to inference. The solution is built using containerization, container orchestration and cloud architecture and built on top of 100% open source foundations.
Production Inference is a critical step on the MDLC to realize the business value for Bloomberg AI applications and the advent of large language models (LLMs) presents new opportunities for expanding NLP capabilities in our products. The inference solution is powered by open source project KServe which is a production ready inference solution for both generative and predictive AI applications. We are poised for enormous user growth this year and have an ambitious roadmap in terms of new features as well as improved user experience. That’s where you come in. As a member of the inference team, you’ll have the opportunity to design and implement scalable, low latency, high throughput model inference solutions in a hybrid cloud environment. We are founding members of the KServe project to standardize ML Inference within the Kubernetes ecosystem. As part of that, we regularly upstream features we develop, present at conferences and collaborate with our peers in the industry. Open source is at the heart of our team. It's not just something we do in our free time, it is how we work.
We’ll trust you to:
Interact with data scientists to understand their production use cases and requirements to advise the next set of GenAI features for the inference platform.
Design solutions for problems such as scalable model deployment, low latency/high throughput inference, GPU resource optimizations and autoscaling.
Automate operation and improve telemetry of the inference platform in our infrastructure stack.
Design solutions for multi-cloud strategy.
You’ll need to be able to:
Innovate and design solutions that keep in mind strict production SLA: low latency/high throughput, multi-tenancy, high availability, reliability across clusters/data centers, etc.
Fix and optimize generative inference application performance.
Provide developer and operational documentation.
Provide performance analysis and capacity planning for clusters.
String communication and collaboration skills, with the ability to work effective with multi-functional teams
Have a passion for providing reliable and scalable infrastructure.
You’ll need to have:
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s