Performance of local Medical Large Language Models as clinical decision support systems under connectivity constraints

Autori

DOI:

https://doi.org/10.5902/2357797597015

Parole chiave:

Clinical decision support systems, Local Large Language Models, Military healthcare support, Offline operational environments, Model quantization

Abstract

In the context of military health support, medical teams frequently operate under extreme conditions, characterized by severe connectivity restrictions, high cognitive pressure, the need for constant mobility, and severe limitations in computational and logistical resources. In this sense, the objective of this study was to investigate the performance of Large Language Models (LLMs), fine-tuned for medical applications, executed locally (offline), considering their application in operational scenarios with connectivity constraints. The experiment followed the classic validation pipeline of Clinical Decision Support Systems (CDSS). The experiment was executed using Python code on two hardware architectures: with a GPU and with a CPU (without a GPU). Eight models were tested using three reference medical benchmarks. The evaluation metrics were: Clinical Accuracy (%), Average Latency per Question (in seconds), and Tokens per Second (TPS). The results show that quantized models in the 8-billion parameter range, especially LlaMA-3.1-8B-Medical (Q4_K_M), have significant potential to act as force multipliers, achieving accuracies of up to 74% in complex medical benchmarks (such as MMLU_Medical and MedMCQA). However, artificial intelligence tools must be integrated into the battlefield exclusively as cognitive assistants under human supervision.

Downloads

I dati di download non sono ancora disponibili.

Biografie autore

Eduardo Borba Neves, Escola de Aperfeiçoamento de Oficiais

Doctor of Biomedical Engineering from the Federal University of Rio de Janeiro; Doctor of Public Health and Environment from the National School of Public Health – ENSP of the Oswaldo Cruz Foundation; Doctor of Notorious Knowledge in Military Education and Culture from the Department of Teaching and Research – DEP of the Brazilian Army; Coordinator, Escola de Aperfeiçoamento de Oficiais, Rio de Janeiro, RJ, Brasil.

Pablo Gustavo Cogo Pochmann, Escola de Aperfeiçoamento de Oficiais

Master of Military Sciences from the Officers' Improvement School – EsAO of the Brazilian Army; Master of Defense Engineering from the Military Engineering Institute – IME of the Brazilian Army; Professor, Escola de Aperfeiçoamento de Oficiais, Rio de Janeiro, RJ, Brasil.

Riferimenti bibliografici

ANGTHONG, C.; RUNGRATTANAWILAI, N.; PUNDEE, C. Artificial intelligence assistance in deciding management strategies for polytrauma and trauma patients. Polish Journal of Surgery, 96, n. Suppl. 1, 2023. 114-117.

DONGARRA, J. et al. Hardware trends impacting floating-point computations in scientific applications. arXiv preprint arXiv:2411.12090, 2024.

HE, K. et al. A survey of large language models for healthcare: from data, technology, and applications to accountability and ethics. Information Fusion, 118, 2025. 102963.

HENDRYCKS, D. et al. Measuring Massive Multitask Language Understanding. arXiv preprint arXiv:2009.03300, 2020.

HOYT, R. E. et al. Evaluating Large Reasoning Model Performance on Complex Medical Scenarios In The MMLU-Pro Benchmark. medRxiv preprint, 2025. 2025-04.

JIN, D. et al. What disease does this patient have? a large-scale open domain question answering dataset from medical exams. Applied Sciences, 11, n. 14, 2021. 6421.

LEONE, R. M. et al. Artificial intelligence in military medicine. Military Medicine, 189, n. 9-10, 2024. 244-248.

MATHAIS, Q. et al. Artificial intelligence for battlefield triage in large-scale combat operations: opportunities, limits, and ethical considerations. Journal of Trauma and Acute Care Surgery, 100, n. 3, 2026. 412-420.

NAZI, Z. A.; PENG, W. Large Language Models in Healthcare and Medical Domain: A Review. Informatics, 11, n. 3, 2024. 57.

NGUYEN, V. A. et al. Quantifying the speed-accuracy trade-off of large language models on oral and maxillofacial surgery multiple-choice questions. Scientific Reports, 15, n. 1, 2025. 40657.

PAL, A.; UMAPATHI, L. K.; SANKARASUBBU, M. Medmcqa: A large-scale multi-subject multi-choice dataset for medical domain question answering. Proceedings of the Conference on Health, Inference, and Learning, 174, 2022. 248-260.

RAVISANKAR, A. V. Combat Casualty Care in the 21st Century: Advances, Challenges, and Evidence-Based Strategies for Armed Forces Medical Services. JSES International, 10, n. 101716, 2026.

RIGGENBACH, Z. W. et al. AI on the Front Lines: A Primer for the Military Health Professional. Military Medicine, 190, n. 9-10, 2025. e1851-e1857.

RIVERA-NICHOLS, T. et al. Investigation and Analysis of Available Chatbot Technologies to Integrate in Multi-Domain Operational, Delayed/Disconnected, Intermittently Connected, Low-Bandwidth Conditions. Military Medicine, 190, n. suppl. 2, 2025. 829-836.

TOUVRON, H. et al. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288, 2023.

WANG, D.; ZHANG, S. Large language models in medical and healthcare fields: applications, advances, and challenges. Artificial Intelligence Review, 57, n. 11, 2024. 299.

WORNOW, M. et al. The shaky foundations of large language models and foundation models for electronic health records. npj Digital Medicine, 6, n. 1, 2023. 135.

XU, Q. et al. Interpretability of Clinical Decision Support Systems Based on Artificial Intelligence from Technological and Medical Perspective: A Systematic Review. Journal of Healthcare Engineering, 2023, 2023. 1-13.

##submission.downloads##

Pubblicato

2026-09-04