PSI - Issue 84

246 Gianluca Bruno et al. / Procedia Structural Integrity 84 (2026) 240–247 normalized difference between the error value at between two successive instants. The minimization of ∆σ drives the agent along more direct trajectories towards the sub-optimal solution. The extraction of the cube root allows to appropriately resize the contribution of ∆σ . Finally, the deviation σ ( ) can be expressed as ( ) = ∑ | − ( ) | =1 (9) In this way, the agent was incentivized to progressively reduce the modal difference, converging towards combinations of parameters compatible with the measurements. The training phase was conducted for a predetermined number of episodes, each started from a random initial configuration of parameters within the permitted limits. The process is governed by two termination criteria: (i) reaching a modal deviation value below a tolerance, i.e., σ ( ) <σ , with σ =0.03 , which represents the tolerance and accuracy value to achieve during the training phase; (ii) exceeding the maximum number of timesteps =100 per episode, to avoid non-convergence. Fig. 4 shows the results of the agent training phase, depicting the trend of cumulative reward in each episode. The moving average of the reward (orange line) stabilized around –15, while the associated deviation remained within the prescribed tolerance, with σ ( ) <0.03 , confirming the stability and convergence of the learning process.

Fig. 3. Results of the training phase for the first case study in terms of cumulative reward per episode with corresponding simple moving average (orange dashed line) At the end of training, the learned policy was tested in a validation phase, considering 1000 episodes and a threshold value σ =0.03 . In this phase, the agent ability to repeatedly converge towards satisfying configurations was assessed. The results obtained from the validation phase of the algorithm are shown in Table 2, close to the real ones, which reports the performance of the algorithm in terms of Root Mean Square Error (RMSE) and Weighted Absolute Percentage Error (WAPE), and the values of and . Table 2. Performance of the proposed algorithm, values of and (first three). Parameter Real Values Proposed method 1 (Hz) 4.7429 4.7697 2 (Hz) 5.7443 5.7668 3 (Hz) 6.6061 6.6882 1 (MPa) 34200 34942 2 (MPa) 33500 33193 3 (MPa) 33450 33735 1 1 1 2 1 1 RMSE --- 322.05 WAPE --- 0.013

Made with FlippingBook flipbook maker