Effective management of diabetes requires patients to actively engage in self-care activities along with medical treatments. Providing on point information to new patients is thus of vital importance. In the last years, we have seen an increasing adoption of Large Language Models (LLMs) to create effective chatbot assistants for diabetes patients. However, despite their diffusion, there is still a limited number of comparative studies about the effectiveness of different LLMs for this task, especially for new inexperienced patients. The goal of this study is to provide a comparison among four state-of-the-art LLMs (ChatGPT4o-mini, Claude3.5-Haiku, Llama3.2-3B-instruct and Mistral-Small-24-02). To this end, we designed D-Care, a chatbot assistant that supports multiple LLMs and tones of voice. Formal experiments, involving both standard metrics and a user study with 40 participants, showed a clear difference in performance between the models, with ChatGPT4o-mini and Claude3.5-Haiku ranking higher than the other two LLMs.

D-Care 2.0: A Multi-tone Multi-LLM Chatbot Assistant for New Diabetes Patients

Nawabi, Awais Khan;Toffanin, Chiara;Dondi, Piercarlo
2026-01-01

Abstract

Effective management of diabetes requires patients to actively engage in self-care activities along with medical treatments. Providing on point information to new patients is thus of vital importance. In the last years, we have seen an increasing adoption of Large Language Models (LLMs) to create effective chatbot assistants for diabetes patients. However, despite their diffusion, there is still a limited number of comparative studies about the effectiveness of different LLMs for this task, especially for new inexperienced patients. The goal of this study is to provide a comparison among four state-of-the-art LLMs (ChatGPT4o-mini, Claude3.5-Haiku, Llama3.2-3B-instruct and Mistral-Small-24-02). To this end, we designed D-Care, a chatbot assistant that supports multiple LLMs and tones of voice. Formal experiments, involving both standard metrics and a user study with 40 participants, showed a clear difference in performance between the models, with ChatGPT4o-mini and Claude3.5-Haiku ranking higher than the other two LLMs.
2026
9783032344588
9783032344595
File in questo prodotto:
Non ci sono file associati a questo prodotto.

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11571/1558616
Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus ND
  • ???jsp.display-item.citation.isi??? ND
social impact