To make high-quality research more accessible and easier to explore.

Fields:
3 results ✕ Clear filters

Online Learning for Dual-Index Policies in Dual-Sourcing Systems

Manufacturing and Service Operations Management 2023
Problem definition: We consider a periodic-review dual-sourcing inventory system with a regular source (lower unit cost but longer lead time) and an expedited source (shorter lead time but higher unit cost) under carried-over supply and backlogged demand. Unlike existing literature, we assume that the firm does not have access to the demand distribution a priori and relies solely on past demand realizations. Even with complete information on the demand distribution, it is well known in the literature that the optimal inventory replenishment policy is complex and state dependent. Therefore, we focus our attention on a class of popular, easy-to-implement, and near-optimal heuristic policies called the dual-index policy. Methodology/results: The performance measure is the regret, defined as the cost difference of any feasible learning algorithm against the full-information optimal dual-index policy. We develop a nonparametric online learning algorithm that admits a regret upper bound of [Formula: see text], which matches the regret lower bound for any feasible learning algorithms up to a logarithmic factor. Our algorithm integrates stochastic bandits and sample average approximation techniques in an innovative way. As part of our regret analysis, we explicitly prove that the underlying Markov chain is ergodic and converges to its steady state exponentially fast via coupling arguments, which could be of independent interest. Managerial implications: Our work provides practitioners with an easy-to-implement, robust, and provably good online decision support system for managing a dual-sourcing inventory system.

Offline Feature-Based Pricing Under Censored Demand: A Causal Inference Approach

Manufacturing and Service Operations Management 2025 27(2), 535-553
Problem definition: We study a feature-based pricing problem with demand censoring in an offline, data-driven setting. In this problem, a firm is endowed with a finite amount of inventory and faces a random demand that is dependent on the offered price and the features (from products, customers, or both). Any unsatisfied demand that exceeds the inventory level is lost and unobservable. The firm does not know the demand function but has access to an offline data set consisting of quadruplets of historical features, inventory, price, and potentially censored sales quantity. Our objective is to use the offline data set to find the optimal feature-based pricing rule so as to maximize the expected profit. Methodology/results: Through the lens of causal inference, we propose a novel data-driven algorithm that is motivated by survival analysis and doubly robust estimation. We derive a finite sample regret bound to justify the proposed offline learning algorithm and prove its robustness. Numerical experiments demonstrate the robust performance of our proposed algorithm in accurately estimating optimal prices on both training and testing data. Managerial implications: The work provides practitioners with an innovative modeling and algorithmic framework for the feature-based pricing problem with demand censoring through the lens of causal inference. Our numerical experiments underscore the value of considering demand censoring in the context of feature-based pricing.

Multiproduct Inventory Systems with Upgrading: Replenishment, Allocation, and Online Learning

Manufacturing and Service Operations Management 2025
Problem definition: We consider the joint optimization of ordering and upgrading decisions in a dynamic multiproduct system over a finite horizon of T periods. In each period, multiple types of demand arrive stochastically and can be satisfied either with supply of the same type or by upgrading to a higher-quality product. The goal is to find an optimal joint replenishment and allocation policy that maximizes total expected profit, both when the firm knows the demand distributions a priori and when the firm must learn them over time. Methodology/results: We first characterize the structure of the clairvoyant optimal joint ordering and allocation policy. Building on this structure, we propose a new online learning algorithm, termed stochastic subgradient descent with perturbed subgradient (SGD-PG for short), and show that it achieves cumulative regret growing on the order of the square root of T, which matches the known lower bound for any online learning method. We further show that SGD-PG can be extended to a nested censored demand setting. In the course of the algorithmic design, we propose a linear programming (LP)-based approach to compute the subgradient and prove that it produces the same output as the perturbed subgradient method. The LP-based method also allows us to extend the results to general upgrading structures. We demonstrate the efficacy of the proposed algorithms in numerical experiments. Managerial implications: This work provides practitioners with the optimal policy for inventory replenishment and allocation in a multiproduct system with upgrading. When the demand distribution is unknown, we propose an easy-to-implement and provably good algorithm for demand learning. In addition, our numerical results quantify the value of optimal upgrading and identify the conditions under which upgrading is most beneficial.