Reinforcement Learning-Based Automated Replenishment System for Optimal Order Quantity Prediction in Grocery Supply Chains
Main Article Content
Abstract
The inventory replenishment of the grocery retail is a complex and stochastic sequential decision with perishable goods, fluctuating demand, promotion times, and competing cost objectives like holding cost, stock-out penalty and wastage. The classical methods involving Economic order quantity (EDOQ) model and the fixed reorder point (ROP) systems are not responsive to the state-dependency of the modern retail system of inventory. The paper presents an Automated Replenishment System (ARS) which is adopted on the premises of Reinforcement Learning (RL) and Markov Decision Processes (MDP), Q-learning, and constrained numerical optimization are used to derive optimal replenishment policies and order quantities on the product-store level. It is experimented on 1.5 million simulation runs using realistic data on the grocery supply chain where replenishment is a decision problem of three actions: (1) Order, (2) Substitute, and (3) Do Nothing and where the rewards are given using composite cost function that incorporates the wastage, stock out cost, and customer satisfaction. The mathematical model involves the Bellman optimality equation to approximate value functions, a constrained quadratic programming layer to determine optimal order quantity, and hierarchical constraint system to address the needs of the labor capacity, physical storage, supplier fulfilment, and merchandising. The results of simulation experiments show statistically significant reductions in aggregate inventory cost, and service level relative to heuristic baseline policies. When the cost of EOQ is reduced by 23.4 percentage points and a random policy is reduced by 31.7 percentage point, the RL-ARS achieves a reduction of 23.4 percentage points and cost-reduction of 31.7 percentage points with an in-stock level of 94.7 and a wastage rate of 3.2.