Effective management of diabetes requires patients to actively engage in self-care activities along with medical treatments. Providing on point information to new patients is thus of vital importance. In the last years, we have seen an increasing adoption of Large Language Models (LLMs) to create effective chatbot assistants for diabetes patients. However, despite their diffusion, there is still a limited number of comparative studies about the effectiveness of different LLMs for this task, especially for new inexperienced patients. The goal of this study is to provide a comparison among four state-of-the-art LLMs (ChatGPT4o-mini, Claude3.5-Haiku, Llama3.2-3B-instruct and Mistral-Small-24-02). To this end, we designed D-Care, a chatbot assistant that supports multiple LLMs and tones of voice. Formal experiments, involving both standard metrics and a user study with 40 participants, showed a clear difference in performance between the models, with ChatGPT4o-mini and Claude3.5-Haiku ranking higher than the other two LLMs.
D-Care 2.0: A Multi-tone Multi-LLM Chatbot Assistant for New Diabetes Patients
Nawabi, Awais Khan;Toffanin, Chiara;Dondi, Piercarlo
2026-01-01
Abstract
Effective management of diabetes requires patients to actively engage in self-care activities along with medical treatments. Providing on point information to new patients is thus of vital importance. In the last years, we have seen an increasing adoption of Large Language Models (LLMs) to create effective chatbot assistants for diabetes patients. However, despite their diffusion, there is still a limited number of comparative studies about the effectiveness of different LLMs for this task, especially for new inexperienced patients. The goal of this study is to provide a comparison among four state-of-the-art LLMs (ChatGPT4o-mini, Claude3.5-Haiku, Llama3.2-3B-instruct and Mistral-Small-24-02). To this end, we designed D-Care, a chatbot assistant that supports multiple LLMs and tones of voice. Formal experiments, involving both standard metrics and a user study with 40 participants, showed a clear difference in performance between the models, with ChatGPT4o-mini and Claude3.5-Haiku ranking higher than the other two LLMs.I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.


