Prediction of BOF endpoint P content by Optuna-RF with K-medoids clustering based on metallurgical parameter constraint
-
Abstract
The accurate prediction of endpoint phosphorus content is very important for controlling the basic oxygen furnace (BOF) steelmaking process. Based on the production data, a random forest (RF) model integrating K-medoids clustering under metallurgical parameter constraints, hyperparameter Optuna optimization, and interpretability analysis (Optuna-RF) is proposed for predicting the endpoint phosphorus content by comparison with four machine learning models: RF, eXtreme Gradient Boosting (XGBoost), Natural Gradient Boosting (NGBoost), and Light Gradient Boosting Machine (LightGBM). The Optuna-RF model uses intelligent sampling based on Optuna’s tree-structured Parzen estimator (TPE) to automatically optimizes the RF hyperparameters with 48% importance of the key parameter max_features contribution, reducing the manual workload of parameter tuning. The measurement error is significantly lower than that of the comparison models, with a root mean square error (RMSE) of 0.00186wt%, which is being 21.1% lower than that of the traditional RF, and a mean absolute error (MAE) of 0.0009686wt%, which is being 83% lower than that of XGBoost. The Optuna-RF model exhibits the smallest error fluctuation range and the most concentrated error distribution (with the highest peak at zero error) in predicting the endpoint phosphorus content. Specifically, it achieves hit rates of 86.55% within the error interval −15, 15 and 97.31% within −25, 25. This demonstrates significantly higher stability and reliability compared to the other four machine learning models. Characteristic importance analysis of the Optuna-RF model using SHapley Additive exPlanations (SHAP) analysis reveals the key process parameters affecting the endpoint phosphorus content, which provides a theoretical basis and data support for the optimization of the steelmaking process and process control.
-
-