Indirect Prompt Injection in Municipal Document-Processing Copilots: Attacks, Defences and Harm
A representative proof-of-concept evaluation across four large language models
DOI:
https://doi.org/10.62161/sauc.v12.6399Keywords:
prompt injection, large language models, AI cybersecurity, case management, eProcurement, e-governmentAbstract
Public administrations are deploying large language model (LLM) assistants that process, summarise, classify and validate citizen-submitted documents. These copilots are exposed to indirect prompt injection: instructions hidden in manipulated documents that reach the model as if they were data.
We develop a municipal aid-procedure copilot and evaluate its robustness across six attack objectives, five delivery vectors, five defences and four LLMs, with 3,000 attack evaluations and 1,000 legitimate evaluations. We combine deterministic detection with a human-validated LLM judge.
Defences reduce attack success, but unevenly across objectives. They largely neutralise imperative attacks, such as request misrouting, while barely affecting summary falsification, revealing a provenance gap: the inability to determine whether an output value originates from an authoritative field or attacker-controlled content. Moreover, attack success and potential harm are decoupled: payment fraud and rule-exfiltration attacks have the greatest potential for harm despite intermediate success rates. We frame the risk within a human-in-the-loop model, where automation bias may amplify harm.
Downloads
Global Statistics ℹ️
|
0
Views
|
0
Downloads
|
|
0
Total
|
|
References
Agencia Española de Protección de Datos. (2025). Política general para el uso de IA generativa en procesos administrativos de la AEPD. https://www.aepd.es/recurso-multimedia/politica-general-interna-para-el-uso-de-ia-generativa-en-procesos
Alon-Barkat, S., & Busuioc, M. (2023). Human–AI interactions in public sector decision making: “Automation bias” and “selective adherence” to algorithmic advice. Journal of Public Administration Research and Theory, 33(1), 153–169. https://doi.org/10.1093/jopart/muac007
Autio, C., Schwartz, R., Dunietz, J., Jain, S., Stanley, M., Tabassi, E., Hall, P., & Roberts, K. (2024). Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile (Number NIST AI 600-1). https://doi.org/10.6028/NIST.AI.600-1
Chen, S., Piet, J., Sitawarin, C., & Wagner, D. (2025). StruQ: Defending against prompt injection with structured queries. 34th USENIX Security Symposium (USENIX Security 25), 2383–2400.
Chiang, C.-H., & Lee, H. (2023). Can large language models be an alternative to human evaluations? Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics, Volume 1, 15607–15631. https://doi.org/10.18653/v1/2023.acl-long.870
Debenedetti, E., Zhang, J., Balunovic, M., Beurer-Kellner, L., Fischer, M., & Tramèr, F. (2024). AgentDojo: A dynamic environment to evaluate prompt injection attacks and defenses for LLM agents. Advances in Neural Information Processing Systems (NeurIPS), 37, Datasets and Benchmarks Track.
European Parliament. (2024). Regulation (EU) 2024/1689 of the European Parliament and of the Council (Artificial Intelligence Act) of 13 June 2024 laying down harmonised rules on artificial intelligence and amending Regulations (EC) No 300/2008, (EU) No 167/2013, (EU) No 168/2013, (EU) 2018/858, (EU) 2018/1139 and (EU) 2019/2144 and Directives 2014/90/EU, (EU) 2016/797 and (EU) 2020/1828. http://data.europa.eu/eli/reg/2024/1689/oj
Gao, T., Yen, H., Yu, J., & Chen, D. (2023). Enabling large language models to generate text with citations. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 6465–6488. https://doi.org/10.18653/v1/2023.emnlp-main.398
Goddard, K., Roudsari, A., & Wyatt, J. C. (2012). Automation bias: A systematic review of frequency, effect mediators, and mitigators. Journal of the American Medical Informatics Association, 19(1), 121–127. https://doi.org/10.1136/amiajnl-2011-000089
Govern de les Illes Balears. (2026). El Govern destina 5,4 millones de euros a incorporar la inteligencia artificial a toda la Administración para ganar agilidad y eficiencia. https://www.caib.es/pidip2front/jsp/es/ficha-convocatoria/el-govern-destina-54-millones-de-euros-a-incorporar-la-inteligencia-artificial-a-toda-la-administracion-para-ganar-agilidad-y-eficiencia
Greshake, K., Abdelnabi, S., Mishra, S., Endres, C., Holz, T., & Fritz, M. (2023). Not what you’ve signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection. https://arxiv.org/abs/2302.12173
Hines, K., Lopez, G., Hall, M., Zarfati, F., Zunger, Y., & Kiciman, E. (2024). Defending against indirect prompt injection attacks with spotlighting. Proceedings of the Conference on Applied Machine Learning in Information Security (CAMLIS 2024), CEUR Workshop Proceedings, 3920, 48–62.
Liu, Y., Jia, Y., Geng, R., Jia, J., & Gong, N. Z. (2024). Formalizing and benchmarking prompt injection attacks and defenses. 33rd USENIX Security Symposium (USENIX Security 24), 1831–1847. https://www.usenix.org/conference/usenixsecurity24/presentation/liu-yupei
Lyell, D., & Coiera, E. (2017). Automation bias and verification complexity: A systematic review. Journal of the American Medical Informatics Association, 24(2), 423–431. https://doi.org/10.1093/jamia/ocw105
Madan, R., & Ashok, M. (2023). AI adoption and diffusion in public administration: A systematic literature review and future research agenda. Government Information Quarterly, 40(1), 101774. https://doi.org/10.1016/j.giq.2022.101774
OWASP Foundation. (2026). OWASP Top 10 for LLM Applications 2026 (Version 2026). https://genai.owasp.org/resource/owasp-genai-llm-top-10-2026/
Red.es. (2024). El sistema de IA en Kit Consulting permite reducir los plazos de comprobación de documentos hasta un 70%. https://www.red.es/es/actualidad/noticias/kit-consulting-incorporacion-ia-justificacion-ayudas
Ruschemeier, H., & Hondrich, L. J. (2024). Automation bias in public administration—An interdisciplinary perspective from law and psychology. Government Information Quarterly, 41(3), 101953. https://doi.org/10.1016/j.giq.2024.101953
Yi, J., Xie, Y., Zhu, B., Kiciman, E., Sun, G., Xie, X., & Wu, F. (2025). Benchmarking and defending against indirect prompt injection attacks on large language models. Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD ’25), 1809–1820. https://doi.org/10.1145/3690624.3709179
Zhan, Q., Fang, R., Panchal, H. S., & Kang, D. (2025). Adaptive attacks break defenses against indirect prompt injection attacks on LLM agents. Findings of the Association for Computational Linguistics: NAACL 2025, 7116–7132. https://doi.org/10.18653/v1/2025.findings-naacl.395
Zhan, Q., Liang, Z., Ying, Z., & Kang, D. (2024). InjecAgent: Benchmarking Indirect Prompt Injections in Tool-Integrated Large Language Model Agents. Proceedings of the Annual Meeting of the Association for Computational Linguistics, 10471–10506. https://doi.org/10.18653/v1/2024.findings-acl.624
Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E. P., Zhang, H., Gonzalez, J. E., & Stoica, I. (2023). Judging LLM-as-a-judge with MT-Bench and Chatbot Arena. Advances in Neural Information Processing Systems, 36, 46595–46623.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Authors retain copyright and transfer to the journal the right of first publication and publishing rights

This work is licensed under a Creative Commons Attribution-NoDerivatives 4.0 International License.
Those authors who publish in this journal accept the following terms:
-
Authors retain copyright.
-
Authors transfer to the journal the right of first publication. The journal also owns the publishing rights.
-
All published contents are governed by an Attribution-NoDerivatives 4.0 International License.
Access the informative version and legal text of the license. By virtue of this, third parties are allowed to use what is published as long as they mention the authorship of the work and the first publication in this journal. If you transform the material, you may not distribute the modified work. -
Authors may make other independent and additional contractual arrangements for non-exclusive distribution of the version of the article published in this journal (e.g., inclusion in an institutional repository or publication in a book) as long as they clearly indicate that the work was first published in this journal.
- Authors are allowed and recommended to publish their work on the Internet (for example on institutional and personal websites), following the publication of, and referencing the journal, as this could lead to constructive exchanges and a more extensive and quick circulation of published works (see The Effect of Open Access).







