A Trend-Augmented Multivariate Regression Model for University Enrollment Forecasting: The Case of Thu Dau Mot University, 2011–2025
Keywords:
university enrollment; multivariate regression; time trend; rolling-origin validation; Diebold–Mariano test; forecast combination; Thu Dau Mot University.Abstract
This study develops a multivariate regression model (M2) to analyze and forecast student enrollment at Thu Dau Mot University from 2011 to 2025, using four explanatory variables: the number of academic programs, two regime-adjustment variables for the COVID-19 and post-COVID periods, and a linear time trend.
The model attains R² = 0.8637 and adjusted R² = 0.8092, with a regression standard error of 668.63 students. In-sample errors are MAE = 412.62, RMSE = 545.94 and MAPE = 10.04%. The trend coefficient is 140.23 students per year (p = 0.0229), and the COVID coefficient is 1,202.23 (p = 0.0145). The maximum VIF is 1.8326, and diagnostic tests do not reject the OLS assumptions.
One-step-ahead expanding-window validation over 2016–2025 yields MAE = 881.01 and MAPE = 19.30%, with a negative bias of ME = −478.54 students.
The model is benchmarked against seventeen alternative methods within the same validation framework, including AICc-selected ARIMA, ARIMAX, three exponential smoothing variants, the Theta method, Ridge regression, Random Forest, Gradient Boosting, a hybrid specification, and forecast combinations. Some important results emerge. First, the ARIMA order choice converges to (0, 1, 0), hence in this case the ARIMA model becomes exactly equivalent to the naïve forecasting method, and no one of the univariate extrapolation methods (ARIMA, exponential smoothing, Theta, naïve with drift) performs better than the naïve oneSecond, even though M2 is one of the leaders, it is not the best approach, since three modifications of M2 (ARIMAX, MAE 822.09; hybrid, 844.03; and Ridge, 880.42) give better results, although not statistically significantly better based on Diebold-Mariano test. Third, an obvious accuracy-bias trade-off emerges: combining M2 and Random Forest reduces bias from -478.54 to -70.33 students at the cost of only 1.6% higher MAE.
A real-time bias-correction experiment yields a negative outcome: errors worsen, and the bias switches signs, showing that the model’s bias should be corrected with asymmetric usage rather than an offset term.
Assuming 51 programs in 2026 and 55 in 2027, the model forecasts 6,754 and 7,064 students, with 95% prediction intervals of [4,684; 8,825] and [4,837; 9,291]. Converted into planning units, the standard error range equates to about 52 full-time faculty positions at a 25:1 student-faculty ratio. We recommend interpreting the forecast as ranges rather than point targets.





