Addressing Budget Allocation and Revenue Allocation in Data Market Environments Using an Adaptive Sampling Algorithm

Jan 1, 2023·
Boxin Zhao
Boxin Zhao
Boxiang Lyu
Boxiang Lyu
,
Raul Castro Fernandez
Mladen Kolar
Mladen Kolar
· 0 min read
Abstract
High-quality machine learning models are dependent on access to high-quality training data. When the data are not already available, it is tedious and costly to obtain them. Data markets help with identifying valuable training data: model consumers pay to train a model, the market uses that budget to identify data and train the model (the budget allocation problem), and finally the market compensates data providers according to their data contribution (revenue allocation problem). For example, a bank could pay the data market to access data from other financial institutions to train a fraud detection model. Compensating data contributors requires understanding data’s contribution to the model; recent efforts to solve this revenue allocation problem based on the Shapley value are inefficient to lead to practical data markets. In this paper, we introduce a new algorithm to solve budget allocation and revenue allocation problems simultaneously in linear time. The new algorithm employs an adaptive sampling process that selects data from those providers who are contributing the most to the model. Better data means that the algorithm accesses those providers more often, and more frequent accesses corresponds to higher compensation. Furthermore, the algorithm can be deployed in both centralized and federated scenarios, boosting its applicability. We provide theoretical guarantees for the algorithm that show the budget is used efficiently and the properties of revenue allocation are similar to Shapley’s. Finally, we conduct an empirical evaluation to show the performance of the algorithm in practical scenarios and when compared to other baselines. Overall, we believe that the new algorithm paves the way for the implementation of practical data markets.
Type
Publication
International Conference on Machine Learning (ICML)
publication
Boxin Zhao
Authors
PhD (2020-2025)

Boxin Zhao was a PhD student in Econometrics and Statistics at University of Chicago, Booth School of Business. His research interests include probabilistic graphical models, functional data analysis and distributed learning, with a focus on developing novel methodologies with both practical applications and theoretical guarantees.

Personal website

Boxiang Lyu
Authors
PhD (2019-2024)

Boxiang Lyu was a PhD student in the Econometrics and Statistics dissertation area at University of Chicago Booth School of Business. Prior to Booth, he obtained a Master of Science in Machine Learning (2019) and a Bachelor of Science in Statistics and Machine Learning (2018) from Carnegie Mellon University.

Personal website

Mladen Kolar
Authors
Professor of Data Sciences and Operations
Mladen Kolar is a Professor of Data Sciences and Operations at the University of Southern California Marshall School of Business and a Visiting Professor of Statistics and Data Science at Mohamed bin Zayed University of Artificial Intelligence. Before joining USC, he was on the faculty of the University of Chicago Booth School of Business. His research is focused on high-dimensional statistical methods, graphical models, varying-coefficient models and data mining, driven by the need to uncover interesting and scientifically meaningful structures from observational data. He is a Fellow of the Institute of Mathematical Statistics.