From Guideline to Gameplay: A Pilot Study on AI Simulation in Medical and Physiotherapy Training
1 More Paper · Full Reading

About this paper
A full audio edition of this paper.
Authors: M. Paukkunen, E. Rintamäki, K. Sinivuori, E. Heikkala, J. Määttä, J. Karppinen
Publication date: 2026
Read the paper: https://doi.org/10.1007/978-3-032-28819-6_42
Source license: Creative Commons Attribution 4.0 International — https://creativecommons.org/licenses/by/4.0/
The authors and publisher do not sponsor or endorse this recording.
Transcript
You’re listening to “From Guideline to Gameplay: A Pilot Study on AI Simulation in Medical and Physiotherapy Training,” by M. Paukkunen and colleagues. Published in 2026.
Maija Paukkunen1envelope symbol, Elina Rintamäki1, Kari Sinivuori1, Eveliina Heikkala1,2, Juhani Määttä1,3, and Jaro Karppinen1 1 Research Unit of Health Sciences and Technology, University of Oulu, Oulu, Finland the email address 2 Wellbeing Services County of Lapland, Oulu, Finland 3 Wellbeing Services County of Northern-Ostrobothnia, Oulu, Finland
Abstract. Background: We designed an educational simulation, ‘MediCampus AI’, using a large language model with Retrieval-Augmented Generation to repli-cate realistic clinical scenarios aligned with evidence-based Current Care Guide-line (CCG). We piloted the simulation to assess its usability for supporting clinical reasoning, patient interaction, and CCG implementation in medical and phys-iotherapy education. Methods: We used Educational Design Research method to develop artificial intelligence (AI)-based simulation prototype. Medical and physiotherapy students (n = 11) were randomized into simulation (SG) and tra-ditional case-based learning groups (TG). User experiences and implementation determinants were measured with study-specific questions, the Determinants of Implementation Behavior Questionnaire (DIBQ-mp) rated using a 7-point Lik-ert scale, and game-generated data.
Results were reported as change in median scores before and after educational interventions. Open-ended questions were analysed using inductive content analysis. Results: A change in median scores in the DIBQ-mp items was observed in “knowledge” (+2 in SG vs. +1 in TG); “skills” (+3 vs. +1.5); “beliefs about capabilities” (+4 vs. +0.5); “beliefs about consequences” (+2 vs. +0.5); “social influences” (+1 vs. +0); “organization” (+3 vs. +1); “behavioral regulation” (+2 vs. +1); “intentions” (+0 vs. +0.5) and “innovation” (+0 vs. +1), “patient” (+1 vs. +1) and “innovation strategy” (+1 vs. +1). Four categories were identified from qualitative data: A) individually supportive game design; B) safe, scalable environment for future needs; C) con-tent enhancements and D) usability/accessibility development needs. Conclusions: Virtual simulation seemed feasible for embedding CCGs into education.
Com-pared with traditional case-based learning, the AI-based simulation showed greater improvements in most implementation determinants and supported competence and guideline-concordant behaviors.
Clinical guidelines, such as the Finnish Current Care Guidelines (CCG), are important for enhancing quality of care. However, the mere existence of a guideline seldom changes treatment practices. Moreover, guidelines rarely explain how to apply recommenda-tions to individual cases through clinical reasoning and shared decision making with the patient, and these skills are still insufficiently practiced in professional training. When developing solutions for practical and complex educational problems, educational design research (EDR) is widely utilized.
Clinical reasoning is a core competence for healthcare professionals and essential for solving patient-specific problems. This process can be described in distinct stages: assessing the patient’s situation, collecting information, processing data, understanding the problem and establishing a diagnosis, setting goals, implementing interventions, evaluating outcomes, and reflecting and learning from the entire process. Learning games can be used effectively to practice clinical reasoning skills in medical education, and they are increasingly utilised to support such learning. A single gameplay session has nearly doubled diagnostic accuracy among medical students and practicing physicians, with sustained improvements at three months.
Learning outcomes achieved with virtual simulation have been shown to transfer to patient cases involving similar problems and increase learner satisfaction. Large language models (LLMs) enable integration of a natural-language interface, dialogue, and interaction into learning games. However, existing educational games have largely overlooked the interpersonal interaction dimension, failing to provide tools for training communication and patient encountering.
This study aimed to design a simulation ‘MediCampus AI’ for teaching clinical reasoning and patient interaction, and to evaluate its usability both in medical and phys-iotherapy training as clinical reasoning is approached differently in these professions. Specifically, we aimed 1) to generate knowledge about user experiences to guide further development, and 2) to identify the determinants (e.g. knowledge, skills, intentions) of implementing CCG to improve educational value.
2 Materials and Methods
We describe an EDR process comprising: (i) orientation and design phase (collaborative designing of new utilizations of emerging artificial intelligence (AI)-based technologies and description of the software platform prototype); (ii) development phase (pilot trial with user groups); and (iii) retrospective phase (evaluation of learning outcomes and CCG implementation determinants with associated strategies to enhance with educational gaming). EDR was chosen for its pragmatic orientation and its aims to generate usable knowledge for improving educational practice. The evaluation was grounded in behav-ioral change theories, specifically Theoretical Domains Framework (TDF) that is used to identify determinants to CCG implementation, linked to Behavior change wheel to guide the selection of appropriate implementation strategies. The description of EDR is presented in Table 1.
2.1 Phase One: Orientation and Design Phase
Design Narrative. A multidisciplinary team including researchers, clinicians (physi-cians and physiotherapists) and software developers started the development project in 2024. The objective was to design an educational simulation using LLM (Open AI GPT-4o) and Retrieval-Augmented Generation (RAG) technology to replicate realistic clinical scenarios aligned with evidence-based recommendations. Following a review of existing simulations, we prioritized authenticity and alignment with clinical guidelines. The goal was to integrate problem-based learning elements into patient cases to create a context-rich learning environment (Supplementary Fig. 1).
The simulation was designed to foster clinical reasoning skills, including assess-ment, diagnostics, and treatment planning focusing specifically on ‘Hand and forearm strain disorders’ CCG, which was used as source data for simulation. Learners interact with diverse virtual patients including varying characteristics and personality types, each accompanied by structured assessment data and individualized treatment plans. The simulation allows learners to navigate cases independently or with the aid of pre-configured options, including the possibility of selecting unnecessary diagnostic tests to prompt reflection on low-value care, and thereby reinforcing critical thinking. The simulation provides case-specific feedback on patient–practitioner communication, diagnostic reasoning, and the appropriateness of treatment plans.
AI generates varied responses on each playthrough, enhancing authenticity and replayability. ‘MediCampus AI’ virtual consultation flow: Patient intake → History taking → Physical examination → Investigations (if needed) → Diagnostics (+ differential diagnosis) → Management plan (patient education, treatment plan) → Feedback & reflection on above mentioned stages and patient encountering.
2.2 Phase Two: Development Phase
Data Collection. Recruitment took place via emails and posters between 5.-18.11.2024 for voluntary medical students at the University of Oulu and physiotherapy students at Oulu University of Applied Sciences. One week prior to the intervention, partici-pants received instructions via email to familiarize themselves with the CCG together with a link to a pre-training Microsoft Forms questionnaire. This included background questions, study-specific items, and the modified Determinants of Implementation Behavior Questionnaire - multiprofessional version (DIBQ-mp).
Sixteen DIBQ-mp items covering 11 TDF domains (Knowledge, Skills, Beliefs about capabilities, Beliefs about consequences, Intentions, Innovation, Organization, Patient, Innovation strategies, Social influences and Behavioral regulation) were rated using a 7-point Likert scale (1 = strongly disagree; 7 = strongly agree) at pre- and post-intervention. Feedback on the learning environments was also collected with three open-ended questions. Participants were randomly allocated to a simulation group (SG) or a traditional case-based learning group (TG). The SG was given two hours to complete simulations of four patient cases in smaller groups.
The TG worked through four corresponding patient case scenarios over 1.5 h, followed by a 30-min feedback session led by an associate professor in physiatry, who gave a review of the correct diagnoses and the examinations that would have been appropriate in each situation. Training was free of charge. Participation was voluntary, written informed consent was obtained, and participants could withdraw at any state of the study. Research permits for this study were obtained from both the University of Oulu and Oulu University of Applied Sciences.
2.3 Phase Three: Retrospective Phase
Regarding quantitative analysis and statistical methods, demographic outcomes were summarized as medians and interquartile ranges. Within-group median changes from pre- to post-intervention were analyzed separately for each group, and P-values reported using the Wilcoxon Signed-Rank test. This non-parametric test was chosen due to non-normal distribution of the data. Analyses were conducted using SPSS statistics, version 29 (IBM Corporation, USA). For qualitive analysis and mixed methods, an inductive content analysis approach was applied to analyze the responses to open-ended questions. The original expressions were extracted and open coded, named and grouped into subcategories based on content similarity and combined into generic categories, which were finally abstracted into main categories.
Drawing on quantitative and qualitative findings, established elements of game-based learning, theory-informed knowledge, we explored how these components could be integrated to enhance the educational use of ‘MediCampus AI’.
3 Results and Discussion 3.1 Quantitative Results
Of the participants, three (27.2%) were fifth-year and six (54.5%) sixth-year medical students, and two (18.2%) were fourth-year physiotherapy students (Supplementary Table 1). Simulation showed advantages in feedback supported learning, enjoyment and motivation, and was perceived as applicable for guideline-aligned training. Motivation to study was similar across groups. Students in SG reported that the visual design, game progression were clear, and rules easy to follow (Supplementary Table 2).
In SG, teams spend an average of 17 min solving a single patient case (range 8–24 min). The average mean overall rating provided by the game’s feedback system was 3.4 out of 5 (range 3–4). The average cost of investigations conducted during the virtual consultation was e4.20 (range: e0–e42). Players arrived at the correct diagnosis in 70% of cases. The game provided feedback on: (i) communication (e.g., clarity and depth, experience of being heard, empathy and reflection); (ii) the accuracy and relevance of the diagnosis, (iii) concreteness and details of the treatment plan, (iv) the comprehensiveness of the patient guidance, and (v) support for treatment adherence. The feedback delivered by the game was appropriate and, in terms of content, consistent with the pre-specified design.
At baseline, the medians across the DIBQ-mp items were similar, with differences of no more than one scale point (data not shown). Eight of the 11 TDF domains showed greater median change in SG compared to the TG. These included: knowledge (+2 vs. +1); skills (+3 vs. +1.5); beliefs about capabilities (+4 vs. +0.5); beliefs about consequences (+2 vs. +0.5); social influences (+1 vs. +0); organization (+3 vs. +1) and behavioral regulation (+2 vs. +1). Conversely, greater median change was observed in the TG, compared to SG, in the following domains: intentions (+0 vs. +0.5) and innovation (+0 vs. +1). Results were equally good between groups for the subscales related to patient and innovation strategy (Table 2).
1 Wilcoxon signed ranks test 3.2 Qualitative Results
Four main categories were identified: AI simulation strengths were A) a game struc-ture that supports the learning process individually, and B) a safe and scalable learn-ing environment aligned with future needs (Supplementary Fig. 1). The development areas consisted of C) content-related enhancements, and D) usability and accessibility improvements (Supplementary Fig. 2). Students experienced interaction with virtual patients as credible and natural, encouraging learning communication skills. Simula-tion elements were perceived to support clinical reasoning with immediate responses to diagnostic questions, and evaluation and treatment of typical patient cases became familiar through repeated gameplay. The simulation was seen as suitable for novice professionals, also broader potential for use across various learner groups was recog-nized.
AI simulation was described as a useful learning environment where practicing was realistic and safe. To better support learning, students recommended expanding the available examinations to allow for greater variety in clinical decision-making, and more complex patient cases. Usability and accessibility improvements included functionality, clarity of the user interface, and guidance.
3.3 Mixed Methods
To increase the ecological validity of edr-based knowledge and the likelihood that edu-cational innovations will be used to transform educational practice, add pedagogical value and successful CCG Implementation, we present a description of future directions in educational simulation development that is based on our quantitative and qualitative results and behavioral change theories (Table 3).
3.4 Discussion
We developed and piloted an AI simulation that creates realistic patient scenarios for training clinical reasoning, decision-making, and patient–practitioner interaction. This is important because medicine requires not only technical proficiency but also human interaction and empathy. In our pilot, the simulation was considered usable for medical and physiotherapy education, delivered personalized and immediate feedback, and supported guideline-aligned practice.
Our findings align with prior research showing that learning games can effectively support clinical reasoning and are generally motivating and acceptable to learners. Recent trial evidence suggesting that even a single gameplay session can substantially improve diagnostic accuracy parallels our observation that simulation can target both cognitive and interpersonal skills. Extending these insights, our study demonstrates how
LLM–driven simulation can provide richer dialogue, context-sensitive feedback, and guideline-concordant prompts.
The simulation environment enables risk-free practice and iterative learning from error, with potential to replace resource-intensive physical setups and to discourage low-value investigations in real clinics. A need for risk-free training is evident to prepare students for real-world settings where enhancing patient safety is a priority. Additionally, AI-generated analytics can identify learning gaps to inform curriculum development, while aggregated de-identified usage data can reveal implementation barriers and guide theory-informed behavior-change strategies. To date, only a small minority of prior studies have based on the selection of game design elements on underlying theories, which has been suggested as a limitation in effectiveness assessments.
We situated this project within EDR pragmatic approach that iterates design, devel-opment, testing, and refinement to generate usable knowledge for improving educational practice. EDR is not a single method; it uses quantitative, qualitative, and mixed methods designs and requires the same transparency and evidential standards as other scientific research. Previous literature has presented 12 strategies to make gamification work in medical education. ‘MediCampus AI’ meets most of them, including making learning fun, improving motivation, using experiential learning, feedback and technology, providing opportunities for repetition, ensuring sustainability and diversity. These elements support its usability also in larger transnational contexts and for the whole musculoskeletal field.
Strengths of this study include a multidisciplinary co-design and a pilot group involv-ing both graduate-stage medical and physiotherapy students. The strength of the simula-tion is explicit mapping to guideline content. In addition, integration of behavior-change theory to link feedback to determinants such as knowledge, skills, beliefs about capabil-ities and consequences, has not been done before in the context of educational gaming. Limitations of this study include small sample size which limits precision and power. We also identified potential ceiling effects on DIBQ-mp domains in the SG. We did not assess long-term retention or transfer to clinical performance, nor did we evaluate cost-effectiveness. However, our primary aim was to evaluate usability.
4 Conclusions
AI-based simulation is a promising tool for teaching and training clinical reasoning and patient interaction while reinforcing guideline-concordant practice. Its updatability, per-sonalized feedback, and potential for safe, scalable, and culturally adapted training make it a viable adjunct to traditional case-based teaching. It can be a feasible way to implement CCGs into education. A greater change was observed in most implementation determi-nants in the SG. It seemed to support students’ competence and guideline-concordant implementation behaviors at least as well as traditional case-based teaching. However, further multi-site evaluations are needed to assess its value at scale.
Acknowledgements. We thank Mikael Karppinen, Eetu Huhtamäki and Jaakko Kujala for your invaluable technical expertise, all study participants, and the Finnish Medical Society Duodecim for permission to use the clinical guideline in this study.
Appendix/Supplementary Material
(See Supplementary Tables 1, 2, 6 and Supplementary Fig. 3).
Supplementary Table 1. Demographic characteristics 1Percentage (frequency); CCG = Current Care Guideline
Supplementary Table 2. Students’ experiences on the learning environments
Supplementary Table 2. (continued) 1 Percentage (frequency); 2 Mean
Supplementary Table 3. Determinants of implementation behavior (facilitators and barriers) for Current Care guideline implementation. Bolded P-values denote for statistically significant within-group difference between pre- and post-intervention.
Supplementary Table 3. (continued)
Supplementary Table 3. (continued) 1 Wilcoxon signed ranks test; NS = non-significant; CCG = Current Care guideline
Supplementary Fig. 1. ‘MediCampus AI’ platform for AI-based education and learning
Supplementary Fig. 2. Strengths of the ‘MediCampus AI’ learning environment
Supplementary Fig. 3. User experience–driven categories for further development of the ‘MediCampus AI’ learning environment
1. Reho, T., Paukkunen, M., Jokihaara, J., et al.: Käypä hoito käyttöön – etä- ja lähikoulutuksella.
yhtä hyviä tuloksia. (Implementing clinical guidelines – equally effective results with remote and in-person training.) Suom Lääkäril 80, e44946 (2025)
2. McKenney, S., Reeves, T.
Conducting Educational Design Research. Routledge, Abingdon
(2018)
3. Koelewijn, G., et al.: Games to support teaching clinical reasoning in health professions.
4. Morgan, D.J., Scherer, L., Pineles, L., et al.: Game-based learning to improve diagnostic.
accuracy: a pilot randomized-controlled trial. Diagnosis 11, 136–141 (2024)
5. Middeke, A., Anders, S., Raupach, T., et al.: Transfer of clinical reasoning trained with.
a serious game to comparable clinical problems: a prospective randomized study. Simul. Healthcare 15, 75–81 (2020)
6. Blanié, A., Amorim, M.A., Meffert, A., et al.: Assessing validity evidence for a serious game.
dedicated to patient clinical deterioration and communication. Adv. Simul. 5, 4 (2020)
7. Alyami, H., Alawami, M., Lyndon, M., et al.: Impact of using a 3D visual metaphor serious.
game to teach history-taking content to medical students: longitudinal mixed methods pilot study. JMIR Ser. Games 7, e13748 (2019)
8. Westera, W.: Why and how serious games can become far more effective: accommodating.
productive learning experiences, learner motivation and the monitoring of learning gains. J. Educ. Technol. Soc. 22, 59–69 (2019)
9. Krishnamurthy, K., Selvaraj, N., Gupta, P., et al.: Benefits of gamification in medical.
10. Cane, J., O’Connor, D., Michie, S.: Validation of the theoretical domains framework for use.
in behaviour change and implementation research. Implement. Sci. 7, 1–17 (2012)
11. Michie, S., Richardson, M., Johnston, M., et al.: The behavior change technique taxonomy.
(v1) of 93 hierarchically clustered techniques: building an international consensus for the reporting of behavior change interventions. Ann. Behav. Med. 46, 81–95 (2013)
12. Käden ja kyynärvarren rasitussairaudet.
Käypä hoito -suositus (Hand and forearm strain dis- orders Current Care Guideline). Suomalaisen Lääkäriseuran Duodecimin ja Suomen Työter-veyslääkäriyhdistyksen asettama työryhmä. Helsinki: Suomalainen Lääkäriseura Duodecim (2024)
13. Paukkunen, M., Ala.Mursula, L., Öberg, B., et al.: Measuring the determinants of implemen-.
tation behavior in multiprofessional rehabilitation. Eur. J. Phys. Rehabil. Med. 59, 488 (2023)
14. Elo, S., Kyngäs, H.: The qualitative content analysis process.
J. Adv. Nurs. 62, 107–115
(2008)
15. Nembhard, I.M., David, G., Ezzeddine, I., et al.: A systematic review of research on empathy.
16. Aster, A., Laupichler, M.C., Zimmer, S., et al.: Game design elements of serious games in.
the education of medical and healthcare professions: a mixed-methods systematic review of underlying theories and teaching effectiveness. Adv. Health Sci. Educ. Theory Pract. 29, 1825–1848 (2024)
17. Shavelson, R.J., Phillips, D.C., Towne, L., et al.: On the science of education design studies.
Educ. Res. 32, 25–28 (2003)
Open Access This chapter is licensed under the terms of the Creative Commons Attribution 4.0 International License (the linked source), which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons license and indicate if changes were made.
The images or other third party material in this chapter are included in the chapter’s Creative Commons license, unless indicated otherwise in a credit line to the material. If material is not included in the chapter’s Creative Commons license and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder.