Desempenho de Large Language Models (LLMs) médicos locais como sistemas de suporte à decisão clínica, sob restrições de conectividade
DOI:
https://doi.org/10.5902/2357797597015Palavras-chave:
Sistemas de suporte à decisão clínica, Modelos de linguagem de grande porte locais, Apoio de saúde militar, Ambientes operacionais desconectados, Quantização de modelosResumo
No contexto do apoio de saúde militar, as equipes médicas frequentemente atuam sob condições extremas, caracterizadas por severas restrições de conectividade, elevada pressão cognitiva, necessidade de mobilidade constante e limitações severas de recursos computacionais e logísticos. Neste sentido, o objetivo deste estudo foi investigar o desempenho de modelos de Large Language Models (LLMs), fine-tuned para aplicações médicas, executados localmente (offline), considerando sua aplicação em cenários operacionais com restrições de conectividade. O experimento seguiu o pipeline clássico de validação de Sistemas de Suporte à Decisão Clínica (Clinical Decision Support Systems - CDSS). O experimento foi executado por meio de um código em Python, em duas estruturas de hardware: com GPU e com CPU (sem GPU). Foram testados oito modelos, utilizando três benchmarks médicos de referência. As métricas de avaliação foram: Acurácia Clínica (%), Latência Média por Questão (em segundos), e Tokens por Segundo (TPS). Os resultados evidenciam que modelos quantizados na faixa de 8 bilhões de parâmetros, em especial o LlaMA-3.1-8B-Medical (Q4_K_M), possuem um potencial significativo para atuar como multiplicadores de força. Ao alcançar acurácias de até 74% em benchmarks médicos complexos (como MMLU_Medical e MedMCQA). Contudo, as ferramentas de inteligência artificial devem ser integradas ao espaço de batalha exclusivamente como assistentes cognitivos sob a supervisão humana.
Downloads
Referências
ANGTHONG, C.; RUNGRATTANAWILAI, N.; PUNDEE, C. Artificial intelligence assistance in deciding management strategies for polytrauma and trauma patients. Polish Journal of Surgery, 96, n. Suppl. 1, 2023. 114-117.
DONGARRA, J. et al. Hardware trends impacting floating-point computations in scientific applications. arXiv preprint arXiv:2411.12090, 2024.
HE, K. et al. A survey of large language models for healthcare: from data, technology, and applications to accountability and ethics. Information Fusion, 118, 2025. 102963.
HENDRYCKS, D. et al. Measuring Massive Multitask Language Understanding. arXiv preprint arXiv:2009.03300, 2020.
HOYT, R. E. et al. Evaluating Large Reasoning Model Performance on Complex Medical Scenarios In The MMLU-Pro Benchmark. medRxiv preprint, 2025. 2025-04.
JIN, D. et al. What disease does this patient have? a large-scale open domain question answering dataset from medical exams. Applied Sciences, 11, n. 14, 2021. 6421.
LEONE, R. M. et al. Artificial intelligence in military medicine. Military Medicine, 189, n. 9-10, 2024. 244-248.
MATHAIS, Q. et al. Artificial intelligence for battlefield triage in large-scale combat operations: opportunities, limits, and ethical considerations. Journal of Trauma and Acute Care Surgery, 100, n. 3, 2026. 412-420.
NAZI, Z. A.; PENG, W. Large Language Models in Healthcare and Medical Domain: A Review. Informatics, 11, n. 3, 2024. 57.
NGUYEN, V. A. et al. Quantifying the speed-accuracy trade-off of large language models on oral and maxillofacial surgery multiple-choice questions. Scientific Reports, 15, n. 1, 2025. 40657.
PAL, A.; UMAPATHI, L. K.; SANKARASUBBU, M. Medmcqa: A large-scale multi-subject multi-choice dataset for medical domain question answering. Proceedings of the Conference on Health, Inference, and Learning, 174, 2022. 248-260.
RAVISANKAR, A. V. Combat Casualty Care in the 21st Century: Advances, Challenges, and Evidence-Based Strategies for Armed Forces Medical Services. JSES International, 10, n. 101716, 2026.
RIGGENBACH, Z. W. et al. AI on the Front Lines: A Primer for the Military Health Professional. Military Medicine, 190, n. 9-10, 2025. e1851-e1857.
RIVERA-NICHOLS, T. et al. Investigation and Analysis of Available Chatbot Technologies to Integrate in Multi-Domain Operational, Delayed/Disconnected, Intermittently Connected, Low-Bandwidth Conditions. Military Medicine, 190, n. suppl. 2, 2025. 829-836.
TOUVRON, H. et al. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288, 2023.
WANG, D.; ZHANG, S. Large language models in medical and healthcare fields: applications, advances, and challenges. Artificial Intelligence Review, 57, n. 11, 2024. 299.
WORNOW, M. et al. The shaky foundations of large language models and foundation models for electronic health records. npj Digital Medicine, 6, n. 1, 2023. 135.
XU, Q. et al. Interpretability of Clinical Decision Support Systems Based on Artificial Intelligence from Technological and Medical Perspective: A Systematic Review. Journal of Healthcare Engineering, 2023, 2023. 1-13.
Downloads
Publicado
Edição
Seção
Licença

Este trabalho está licenciado sob uma licença Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License.


