PSI - Issue 84
244 Gianluca Bruno et al. / Procedia Structural Integrity 84 (2026) 240–247 the agent receives the observations vector ( ) = { ( ) , ( ) } and perform actions according to the vector ( ) = ( ( ) ) = [ ( 1 ( ) ), … , ( ( ) ), … , ( ( ) )] to modify individual parameters ( ) of the ROM. For each timestep , the aeent receives an instant value of reward ( ) , computed through a reward function defined as ( ) = ( ( ) , ( ) , ) . In this way, the agent is driven to explore the parameter space for searching the configurations that progressively reduce the model updating error. Each episode terminates when the simulated target variables ( ) converge towards the reference values , within a predefined tolerance set by the analyst. If convergence is not achieved after all iterations within an episode, the process is stopped and the episode is marked as unsuccessful. The interaction between the agent and the environment is governed by a policy function ( ) , which the agent progressively learns and updates throughout the episodes. The last two steps of the proposed framework consist of agent training, validation and model selection. The training phase is controlled by three main parameters, defined by the analyst: (a) the maximum number of iterations or timesteps allowed in each episode to achieve a solution; (b) the tolerance value σ , which states when the solution can be considered acceptable; (c) the maximum number of total iterations allowed for the agent to learn a stable policy. At the end of training, the update is considered effective if the agent converges stably towards a solution that reduces the modal error below a tolerance σ , obtaining the solution to the model updating problem. During the validation phase, the agent is tested on a finite number of episodes , but without updating the policy (which remains the same as the optimal learning policy ∗ ( ) ). If at least 90% of the episodes converge towards the solution (error between simulated target variables ( ) and real target variables below the tolerance). If the validation works, the model selection can be performed and the optimal parameters ̂ can be statistically defined, by selecting the mode of the probability distributions of the individual parameters. 4. Application to a real-life RC bridge The case study used to test the proposed framework is an existing three-span prestressed RC bridge. For reasons of confidentiality, no detailed information is provided on the location and name of the structure. From the geometrical point of view, the deck has a total longitudinal length of 87 m, divided into three spans of 27 m, 33 m and 27 m, and a transversal width of approximately 13 m. Each span consists of four main reinforced concrete box girders beams, 1.80 m high, arranged longitudinally and connected by five end and intermediate transverse beams. Fig. 2 reports some images of the bridge, illustrating the configuration of the deck and the arrangement of the box girders. The deck is supported by two frame piers, constituted by three circular RC columns with a diameter of 1.30 m and a pier cap 1.50 m high, and two RC abutments. The total height of the substructures is approximately 6 m. From a static point of view, the bridge presents three simply supported beams, with two expansion joints located at the abutments. However, a continuous 30 cm thick cast-in-place concrete slab connects the longitudinal beams of the different spans, introducing a certain structural continuity between the deck segments.
Fig. 2. Images of the case study bridge.
Made with FlippingBook flipbook maker