PSI - Issue 84

Gianluca Bruno et al. / Procedia Structural Integrity 84 (2026) 240–247

241

1. Introduction Structural health monitoring (SHM) of existing bridges is a fundamental task for infrastructure resilience, since it allows to ensure safety and functionality while accounting for ageing and environmental effects. Continuous monitoring systems provide valuable information on the dynamic response of bridges, suggesting variations in structural behavior, due to deterioration or damage. Still, data collected can be used to create numerical models (i.e., digital twin) to reproduce real-life behavior of the structure and drive decision-making strategies for assessment and maintenance (Giagopoulos et al. 2019, Gentile et al. 2015). The reliability of digital twins depends on an adequate model updating process, in which uncertain parameters (e.g., materials, boundary conditions, restraints) are calibrated against experimental data (Bruno et al. 2025, Ierimonti et al. 2020, Ruggieri et al. 2024). In typical vibration-based applications, modal updating is driven by modal quantities, such as natural frequencies and modal shapes, usually identified by operational modal analysis (Brincker et al. 2001). Even though it seems simple, when the structure is complex and the number of uncertain parameters increases, the process of model updating becomes high-dimensional. In such cases, the trade-off between the updating strategy and model fidelity represents the key aspect of the process, balancing the accuracy of the solution and the related computational cost. Regarding the updating strategies, two main methodologies can be identified: (a) deterministic; (b) probabilistic (Simoen et al. 2015). The first ones are characterized by an optimization problem, in which the error between numerical and experimental data is minimized by iteratively adjusting uncertain parameters. The second ones, mainly based on Bayesian approaches, provide a more rigorous framework for managing uncertainty by treating parameters as random variables and updating their probability distributions according to observed data. Concerning model fidelity, two extremes can be mentioned: (i) full-order models (FOMs); (ii) reduced-order models (ROMs) (Fresca and Manzoni 2022, Vlachas et al. 2021). The first ones enable a detailed representation of local nonlinearities and damage mechanisms, while increasing the computational costs. The second ones efficiently capture global structural behavior, exploiting mathematical spatial and physical reduction, lightening then the related analysis effort. To overcome critical issues associated with accuracy and computational efficiency of the solution, hybrid approaches could be exploited, by combining light numerical models as the ROMs and by selecting as model updating strategy a data-driven approach, which automatically performs the model selection, while tuning uncertain parameters. In this perspective, this paper proposes an advanced model updating framework, which combines the computational efficiency of ROMs with the capability of a deep reinforcement learning (DRL)-based agent in calibrating uncertain parameters of the model. The main advantages of the proposed tool are related to the possibility of managing high dimensional spaces (i.e., increasing number of uncertain parameters), by reducing training and validation time while retaining accuracy of results. After describing the procedure, the proposed approach was tested on a real-life reinforced concrete (RC) bridge. 2. Background 2.1. Fundamentals of deep reinforcement learning DRL is a branch of machine learning in which an autonomous agent is trained to learn a sub-optimal decision strategy through direct interaction with a dynamic environment. Formally, the problem is described as a Markov Decision Process defined by the tuple ( , , ) : • represents the space of states of the environment. • represents the space of the possible actions. • represents the reward function. At each timestep , the agent observes the state ( ) ∈ and performs an action ( ) ∈ according to a policy (s) ∈Π , receiving a observations ( ) ∈Ω (where Ω represents the range of possible observations), and reward ( ) ∈ that quantifies the goodness of the action undertaken with respect to the final objective (Bellman 1957). The process terminates when the solution is reached, defining the episode , i.e., the set of timesteps necessary. The iterative process defining the interaction between agent and environment is shown in Fig. 1.

Made with FlippingBook flipbook maker