Agentes de IA auditables de lazo cerrado para gemelos digitales de puentes urbanos

Validación sintética seguida de una implementación a escala municipal en Utsunomiya

Autores/as

  • Takayuki Shinohara AIST

DOI:

https://doi.org/10.62161/sauc.v12.6387

Palabras clave:

IA agéntica, Gemelos digitales urbanos, Ciudades inteligentes, Mantenimiento de puentes, Resiliencia urbana, Planificación multiobjetivo, IA auditable

Resumen

El mantenimiento de los puentes urbanos es un problema de gobernanza de la ciudad inteligente que implica evidencia incierta, presupuestos restringidos y acciones críticas para la seguridad. Presentamos una arquitectura de inteligencia artificial auditable de lazo cerrado que separa las observaciones, los estados inferidos, las alternativas, las decisiones y las intervenciones. Los mecanismos se validan en SynthTown, un municipio de verdad de referencia con 50 puentes y 842 componentes, mediante la predicción de estados ocultos, la planificación multiobjetivo, el cribado por parte de los actores interesados, la recalibración por retroalimentación y salvaguardas de verificación-reparación. Una implementación con datos públicos abarca 1733 puentes de Utsunomiya e integra 1592 puntos de radar de apertura sintética interferométrico de dispersores persistentes, registros de inspección, modelos de mantenimiento tridimensionales y planes conscientes de la capacidad. La verificación eliminó las salidas inseguras en el banco de pruebas sintético

Descargas

Los datos de descargas todavía no están disponibles.

Estadísticas globales ℹ️

Totales acumulados desde su publicación
0
Visualizaciones
0
Descargas
0
Total
Descargas por formato:
PDF 0 PDF (English) 0

Citas

Amershi, S., Weld, D., Vorvoreanu, M., Fourney, A., Nushi, B., Collisson, P., Suh, J., Iqbal, S.,

Bennett, P. N., Inkpen, K., Teevan, J., Kikin-Gil, R., & Horvitz, E. (2019). Guidelines for human-AI interaction. Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems, Article 3, 1-13. https://doi.org/10.1145/3290605.3300233

Andriotis, C. P., & Papakonstantinou, K. G. (2021). Deep reinforcement learning driven inspection and maintenance planning under incomplete information and constraints. Reliability Engineering & System Safety, 212, 107551. https://doi.org/10.1016/j.ress.2021.107551

Benjamini, Y., & Hochberg, Y. (1995). Controlling the false discovery rate: A practical and powerful approach to multiple testing. Journal of the Royal Statistical Society: Series B, 57(1), 289-300. https://doi.org/10.1111/j.2517-6161.1995.tb02031.x

Bocchini, P., & Frangopol, D. M. (2011). A probabilistic computational framework for bridge network optimal maintenance scheduling. Reliability Engineering & System Safety, 96(2), 332-349. https://doi.org/10.1016/j.ress.2010.09.001

Boje, C., Guerriero, A., Kubicki, S., & Rezgui, Y. (2020). Towards a semantic construction digital twin: Directions for future research. Automation in Construction, 114, 103179. https://doi.org/10.1016/j.autcon.2020.103179

Deb, K., Pratap, A., Agarwal, S., & Meyarivan, T. (2002). A fast and elitist multi-objective genetic algorithm: NSGA-II. IEEE Transactions on Evolutionary Computation, 6(2), 182-197. https://doi.org/10.1109/4235.996017

Debenedetti, E., Zhang, J., Balunovic, M., Beurer-Kellner, L., Fischer, M., & Tramer, F. (2024). AgentDojo: A dynamic environment to evaluate prompt injection attacks and defenses for LLM agents. arXiv preprint arXiv:2406.13352. https://arxiv.org/abs/2406.13352

Dembski, F., Wossner, U., Letzgus, M., Ruddat, M., & Yamu, C. (2020). Urban digital twins for smart cities and citizens: The case study of Herrenberg, Germany. Sustainability, 12(6), 2307. https://doi.org/10.3390/su12062307

Dwork, C., Feldman, V., Hardt, M., Pitassi, T., Reingold, O., & Roth, A. (2015). The reusable holdout: Preserving validity in adaptive data analysis. Science, 349(6248), 636-638. https://doi.org/10.1126/science.aaa9375

Frangopol, D. M., Dong, Y., & Sabatino, S. (2017). Bridge life-cycle performance and cost: Analysis, prediction, optimisation and decision-making. Structure and Infrastructure Engineering, 13(10), 1239-1257. https://doi.org/10.1080/15732479.2016.1267772

Frangopol, D. M., Kong, J. S., & Gharaibeh, E. S. (2001). Reliability-based life-cycle management of highway bridges. Journal of Computing in Civil Engineering, 15(1), 27-34. https://doi.org/10.1061/(ASCE)0887-3801(2001)15:1(27)

Frangopol, D. M., & Liu, M. (2007). Maintenance and management of civil infrastructure based on condition, safety, optimization, and life-cycle cost. Structure and Infrastructure Engineering, 3(1), 29-41. https://doi.org/10.1080/15732470500253164

Gebru, T., Morgenstern, J., Vecchione, B., Vaughan, J. W., Wallach, H., Daume III, H., & Crawford, K. (2021). Datasheets for datasets. Communications of the ACM, 64(12), 86-92. https://doi.org/10.1145/3458723

Geifman, Y., & El-Yaniv, R. (2017). Selective classification for deep neural networks. Advances in Neural Information Processing Systems, 30, 4878-4887. https://proceedings.neurips.cc/paper/2017/hash/4a8423d5e91fda00bb7e46540e2b0cf1-Abstract.html

Guo, C., Pleiss, G., Sun, Y., & Weinberger, K. Q. (2017). On calibration of modern neural networks. Proceedings of the 34th International Conference on Machine Learning, 70, 1321-1330. https://proceedings.mlr.press/v70/guo17a.html

Guo, Z., Cheng, S., Wang, H., Liang, S., Qin, Y., Li, P., Liu, Z., Sun, M., & Liu, Y. (2024). StableToolBench: Towards stable large-scale benchmarking on tool learning of large language models. arXiv preprint arXiv:2403.07714. https://arxiv.org/abs/2403.07714

Halfawy, M. R., Dridi, L., & Baker, S. (2009). A multi-objective optimization decision support model for renewal planning of sewer networks. Journal of Water Management Modeling, R235-09. https://doi.org/10.14796/JWMM.R235-09

Howard, R. A. (1966). Information value theory. IEEE Transactions on Systems Science and Cybernetics, 2(1), 22-26. https://doi.org/10.1109/TSSC.1966.300074

Ji, Z., Lee, N., Frieske, R., Yu, T., Su, D., Xu, Y., Ishii, E., Bang, Y., Chen, D., Dai, W., Madotto, A., & Fung, P. (2023). Survey of hallucination in natural language generation. ACM Computing Surveys, 55(12), 1-38. https://doi.org/10.1145/3571730

Jimenez, C. E., Yang, J., Wettig, A., Yao, S., Pei, K., Press, O., & Narasimhan, K. (2023). SWE-bench: Can language models resolve real-world GitHub issues? arXiv preprint arXiv:2310.06770. https://arxiv.org/abs/2310.06770

Kaelbling, L. P., Littman, M. L., & Cassandra, A. R. (1998). Planning and acting in partially observable stochastic domains. Artificial Intelligence, 101(1-2), 99-134. https://doi.org/10.1016/S0004-3702(98)00023-X

Kritzinger, W., Karner, M., Traar, G., Henjes, J., & Sihn, W. (2018). Digital twin in manufacturing: A categorical literature review and classification. IFAC-PapersOnLine, 51(11), 1016-1022. https://doi.org/10.1016/j.ifacol.2018.08.474

Lei, X., Xia, Y., Deng, L., & Sun, L. (2022). A deep reinforcement learning framework for life-cycle maintenance planning of regional deteriorating bridges using inspection data. Structural and Multidisciplinary Optimization, 65. https://doi.org/10.1007/s00158-022-03210-3

Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Kuttler, H., Lewis, M., Yih, W., Rocktaschel, T., Riedel, S., & Kiela, D. (2020). Retrieval-augmented generation for knowledge-intensive NLP tasks. Advances in Neural Information Processing Systems, 33, 9459-9474. https://arxiv.org/abs/2005.11401

Lin, S., Hilton, J., & Evans, O. (2022). Teaching models to express their uncertainty in words. Transactions on Machine Learning Research. https://arxiv.org/abs/2205.14334

Liu, K., & El-Gohary, N. (2021). Semantic neural network ensemble for automated dependency relation extraction from bridge inspection reports. Journal of Computing in Civil Engineering, 35(4), 04021007. https://doi.org/10.1061/(ASCE)CP.1943-5487.0000961

Liu, P., Xiong, R., & Tang, P. (2022). Mining observation and cognitive behavior process patterns of bridge inspectors. arXiv preprint arXiv:2205.09257. https://arxiv.org/abs/2205.09257

Liu, X., Yu, H., Zhang, H., Xu, Y., Lei, X., Lai, H., Gu, Y., Ding, H., Men, K., Yang, K., Zhang, S., Deng, X., Zeng, A., Du, Z., Zhang, C., Shen, S., Zhang, T., Su, Y., Sun, H., ... Tang, J. (2024). AgentBench: Evaluating LLMs as agents. International Conference on Learning Representations. https://arxiv.org/abs/2308.03688

Lu, J., Holleis, T., Zhang, Y., Aumayer, B., Nan, F., Bai, F., Ma, S., Ma, S., Li, M., Yin, G., Wang, Z., & Pang, R. (2024). ToolSandbox: A stateful, conversational, interactive evaluation benchmark for LLM tool use capabilities. arXiv preprint arXiv:2408.04682. https://arxiv.org/abs/2408.04682

Ma, C., Zhang, J., Zhu, Z., Yang, C., Yang, Y., Jin, Y., Lan, Z., Kong, L., & He, J. (2024). AgentBoard: An analytical evaluation board of multi-turn LLM agents. Advances in Neural Information Processing Systems. https://arxiv.org/abs/2401.13178

Madanat, S., & Ben-Akiva, M. (1994). Optimal inspection and repair policies for infrastructure facilities. Transportation Science, 28(1), 55-62. https://doi.org/10.1287/trsc.28.1.55

Meerow, S., Newell, J. P., & Stults, M. (2016). Defining urban resilience: A review. Landscape and Urban Planning, 147, 38-49. https://doi.org/10.1016/j.landurbplan.2015.11.011

Mialon, G., Fourrier, C., Swift, C., Wolf, T., LeCun, Y., & Scialom, T. (2023). GAIA: A benchmark for general AI assistants. arXiv preprint arXiv:2311.12983. https://arxiv.org/abs/2311.12983

Ministry of Land, Infrastructure, Transport and Tourism, Japan. (2024). White paper on infrastructure management 2024. https://www.mlit.go.jp/statistics/content/001855598.pdf

Morcous, G. (2006). Performance prediction of bridge deck systems using Markov chains. Journal of Performance of Constructed Facilities, 20(2), 146-155. https://doi.org/10.1061/(ASCE)0887-3828(2006)20:2(146)

Moreau, L., & Missier, P. (Eds.). (2013). PROV-DM: The PROV data model. W3C Recommendation. https://www.w3.org/TR/prov-dm/

NASA Earth Science and Technology Office. (n.d.). Technology readiness levels. https://esto.nasa.gov/trl/

Orcesi, A. D., & Frangopol, D. M. (2010). Optimization of bridge management under budget constraints: Role of structural health monitoring. Transportation Research Record: Journal of the Transportation Research Board, 2202(1), 148-155. https://doi.org/10.3141/2202-18

Page, E. S. (1954). Continuous inspection schemes. Biometrika, 41(1-2), 100-115. https://doi.org/10.1093/biomet/41.1-2.100

Papakonstantinou, K. G., & Shinozuka, M. (2014a). Planning structural inspection and maintenance policies via dynamic programming and Markov processes. Part I: Theory. Reliability Engineering & System Safety, 130, 202-213. https://doi.org/10.1016/j.ress.2014.04.005

Papakonstantinou, K. G., & Shinozuka, M. (2014b). Planning structural inspection and maintenance policies via dynamic programming and Markov processes. Part II: POMDP implementation. Reliability Engineering & System Safety, 130, 214-224. https://doi.org/10.1016/j.ress.2014.04.006

Pastore, T., Mariniello, G., & Asprone, D. (2024). A simheuristic approach to scheduling sustainable and reliable maintenance for bridge infrastructure. Mathematics, 12(21), 3420. https://doi.org/10.3390/math12213420

Pereira, G. V., Parycek, P., Falco, E., & Kleinhans, R. (2018). Smart governance in the context of smart cities: A literature review. Information Polity, 23(2), 143-162. https://doi.org/10.3233/IP-170067

Robelin, C.-A., & Madanat, S. M. (2007). History-dependent bridge deck maintenance and replacement optimization with Markov decision processes. Journal of Infrastructure Systems, 13(3), 195-201. https://doi.org/10.1061/(ASCE)1076-0342(2007)13:3(195)

Ruan, Y., Dong, H., Wang, A., Pitis, S., Zhou, Y., Ba, J., Dubois, Y., Maddison, C. J., & Hashimoto, T. (2024). Identifying the risks of LM agents with an LM-emulated sandbox. International Conference on Learning Representations. https://arxiv.org/abs/2309.15817

Shinn, N., Cassano, F., Berman, E., Gopinath, A., Narasimhan, K., & Yao, S. (2023). Reflexion: Language agents with verbal reinforcement learning. Advances in Neural Information Processing Systems, 36. https://arxiv.org/abs/2303.11366

Su, S., Zhong, R. Y., Jiang, Y., Song, J., Fu, Y., & Cao, H. (2023). Digital twin and its potential applications in construction industry: State-of-art review and a conceptual framework. Advanced Engineering Informatics, 58, 102030. https://doi.org/10.1016/j.aei.2023.102030

Thompson, P. D., Small, E. P., Johnson, M., & Marshall, A. R. (1998). The Pontis bridge management system. Structural Engineering International, 8(4), 303-308. https://doi.org/10.2749/101686698780488758

Torres-Machi, C., Yepes, V., & Pellicer, E. (2020). Markov-based deterioration modelling for long-term bridge management: A critical review and research needs. Engineering Structures, 206, 110096. https://doi.org/10.1016/j.engstruct.2020.110096

Wirtz, B. W., Weyerer, J. C., & Geyer, C. (2019). Artificial intelligence and the public sector-Applications and challenges. International Journal of Public Administration, 42(7), 596-615. https://doi.org/10.1080/01900692.2018.1498103

Wu, C., Wu, P., Wang, J., Jiang, R., Chen, M., & Wang, X. (2021). Critical review of data-driven decision-making in bridge operation and maintenance. Structure and Infrastructure Engineering, 18(1), 47-70. https://doi.org/10.1080/15732479.2020.1833946

Yao, S., Shinn, N., Razavi, P., & Narasimhan, K. (2024). Tau-bench: A benchmark for tool-agent-user interaction in real-world domains. arXiv preprint arXiv:2406.12045. https://arxiv.org/abs/2406.12045

Yao, S., Zhao, J., Yu, D., Du, N., Shafran, I., Narasimhan, K., & Cao, Y. (2023). ReAct: Synergizing reasoning and acting in language models. International Conference on Learning Representations. https://arxiv.org/abs/2210.03629

Zhang, C., Karim, M. M., & Qin, R. (2022). A multitask deep learning model for parsing bridge elements and segmenting defects in bridge inspection images. arXiv preprint arXiv:2209.02190. https://arxiv.org/abs/2209.02190

Zhang, C., Lei, X., Xia, Y., & Sun, L. (2024). Automatic bridge inspection database construction through hybrid information extraction and large language models. Developments in the Built Environment, 20, 100549. https://doi.org/10.1016/j.dibe.2023.100549

Zhang, Z., Cui, S., Lu, Y., Zhou, J., Yang, J., Wang, H., & Huang, M. (2024). Agent-SafetyBench: Evaluating the safety of LLM agents. arXiv preprint arXiv:2412.14470. https://arxiv.org/abs/2412.14470

Zhou, S., Xu, F. F., Zhu, H., Zhou, X., Lo, R., Sridhar, A., Cheng, X., Ou, T., Bisk, Y., Fried, D., Alon, U., & Neubig, G. (2023). WebArena: A realistic web environment for building autonomous agents. arXiv preprint arXiv:2307.13854. https://arxiv.org/abs/2307.13854

Publicado

2026-09-03

Cómo citar

Shinohara, T. (2026). Agentes de IA auditables de lazo cerrado para gemelos digitales de puentes urbanos: Validación sintética seguida de una implementación a escala municipal en Utsunomiya. Street Art & Urban Creativity, 12(5), 345–369. https://doi.org/10.62161/sauc.v12.6387

Número

Sección

Artículos de investigación