Physical therapists’ perspectives on a large language model-powered knowledge translation tool for guideline adherence: A qualitative focus group study
1 More Paper · Full Reading

About this paper
A full audio edition of this paper.
Authors: D. Rosen, D. Zwanzig, B. Vogel, M. Erhart, N.L. Reiter
Publication date: 2026
Read the paper: https://doi.org/10.1080/09593985.2025.2606058
Source license: Creative Commons Attribution 4.0 International — https://creativecommons.org/licenses/by/4.0/
The authors and publisher do not sponsor or endorse this recording.
Transcript
You’re listening to “Physical therapists’ perspectives on a large language model-powered knowledge translation tool for guideline adherence: A qualitative focus group study,” by D. Rosen and colleagues. Published in 2026.
Physiotherapy Theory and Practice
An International Journal of Physical Therapy
ISSN: 0959-3985 (Print) 1532-5040 (Online) Journal homepage: the linked source
Physical therapists’ perspectives on a large language model-powered knowledge translation tool for guideline adherence: A qualitative focus group study
Diane Rosen, Dorian Zwanzig, Barbara Vogel, Michael Erhart & Nils Lennart Reiter
To cite this article: Diane Rosen, Dorian Zwanzig, Barbara Vogel, Michael Erhart & Nils Lennart Reiter (2026) Physical therapists’ perspectives on a large language model-powered knowledge translation tool for guideline adherence: A qualitative focus group study, Physiotherapy Theory and Practice, 42:7, 967-980, DOI: 10.1080/09593985.2025.2606058
QUALITATIVE RESEARCH REPORT
Physical therapists’ perspectives on a large language model-powered knowledge translation tool for guideline adherence: A qualitative focus group study
Diane Rosen PT M.Sca,b, Dorian Zwanzig M.Scc, Barbara Vogeld, Michael Erhart Dr.a, and Nils Lennart Reiter PT B.Sc.a,e aDepartment of Health and Education, Alice Salomon University of Applied Sciences Berlin, Berlin, Germany; bBerlin School of Public Health, Charité, Berlin, Germany; cDepartment of Informatics, Communication and Economy, Hochschule für Technik und Wirtschaft Berlin, Universitiy of Applied Sciences, Berlin, Germany; dDepartment for Orthopaedics and Sports Orthopaedics, Physiotherapy, TUM University Hospital Rechts der Isar, Munich, Germany; eSchool of Health Professions, Department of Physiotherapy, Bern University of Applied Sciences, Bern, Switzerland
ABSTRACT.
ARTICLE HISTORY
Background: Clinical practice guidelines support healthcare professionals in making evidence-based decisions, yet guideline adherence among physical therapists remains inconsistent. To address this gap, a prototype digital knowledge translation tool powered by a large language model (LLM) was developed, with content based on two exemplary high-quality German national guidelines. Objective: To explore the experiences of German physical therapists using the tool, assess their perspectives on its utilization in clinical practice, and compare perceptions between outpatient and inpatient settings.
Methods: Six focus group interviews were conducted: three in a university hospital inpatient setting and three in outpatient physical therapy practices. Discussions were analyzed using qualitative content analysis with inductive and deductive coding.
Results: Twenty physical therapists (11 inpatient, 9 outpatient) participated. Overall experiences were positive, though prolonged response times were criticized. Utilization was thought to depend on time availability and workplace digitization. The tool’s potential assisting with clinical questions was highlighted. No considerable differences in experiences across settings were noted. Inpatient therapists envisioned using the tool between sessions for personal knowledge enhancement, whereas outpatient therapists anticipated utilization during sessions for patient education. Conclusion: LLM-based knowledge translation tools may contribute to improving guideline adherence among physical therapists. Successful implementation requires assessment of digital infrastructure, relevance to clinical needs, and users’ digital literacy.
Further research should evaluate the quality of LLM-generated summaries to ensure validity and trustworthiness, and optimize the tools’ usability regarding speed and content. Development should also prompt ethical considerations about their role in clinical decision-making and patient care.
Background.
Implementing scientific knowledge into clinical practice is usually a slow process, with significant delays before new evidence is widely adopted (Finney Rutten, Ridgeway, and Griffin, 2024; Hanney et al, 2015; Morris, Wooding, and Grant, 2011). Knowledge translation strategies, such as evidence-based clinical practice guidelines (CPGs) aim to accelerate this process and overcome barriers to implementation (Bhuiya et al, 2024; Bland et al, 2024; Straus et al, 2010).
CPGs are systematically developed statements to assist practitioners and patients in making informed healthcare decisions for specific clinical circumstances (Institute of Medicine (US), 1990; 2011). Research suggests that adherence to guideline-based physical therapy
(PT) improves patient outcomes, reduces healthcare utilization, and yields significant cost savings. Despite their importance, adherence to CPGs among physical therapists (PTs) has been reported as inconsistent and often inadequate.
In Germany, CPGs are publicly available as extensive PDF files, mostly exceeding 100 pages (Association of the Scientific Medical Societies in Germany (AWMF), 2025). Although CPGs were primarily designed for physicians and lacked specific recommendations for allied healthcare professions, guideline commissions are becoming more interdisciplinary, increasingly incorporating content relevant to PT. In general, German PTs are interested in utilizing CPGs, but several barriers to guideline-adherent practice persist. These factors include awareness of CPGs, their format, limited time to read them, and insufficient scientific knowledge to interpret them.
Most German PTs work in private practices, encountering challenges such as tight scheduling, limited time between sessions, and a variety of health topics, while a smaller proportion work in hospital inpatient settings, often specialized in specific health areas but having to collaborate within interprofessional teams.
Guideline implementation strategies vary considerably across all professions, with a predominance of multifaceted interventions tailored to specific healthcare settings. International literature on implementation strategies for guideline adherence among PTs has primarily focused on barriers related to practitioners’ lack of knowledge, ability, or awareness of CPGs. However, strategies that address barriers related to CPG format and time constraints by improving accessibility to CPG content have not been extensively reported, presenting an opportunity to enhance evidence translation into PT practice.
Meanwhile, the evolving field of large language models (LLMs) has emerged as a promising solution for knowledge translation across various domains, particularly in healthcare education. LLMs are AI systems based on deep learning that use transformer architecture and are trained on large text collections to recognize patterns in language and generate human-like responses. These models offer substantial possibilities for enhancing learning and information dissemination. Recent publications describe their use in PT for a wide range of purposes, including motion analysis, activity monitoring, patient assessment, and clinical decision support, and focus on integrating these technologies into practice for specific diseases or clinical contexts.
In contrast, the “Innovation in Evidence-Based Practice” (In-EbP) project explores how recent technological advancements could enhance the accessibility of CPGs in PT practice. The project followed a Design Science Research approach combined with User-Centered Design principles, and was conducted in two iterative development phases. In the first phase, two competing prototypes were built from scratch and optimized continuously through iterative feedback from eleven outpatient PTs. Then, in a second phase, a combined prototype was developed, integrating features of both initial approaches. This final prototype functions as a digital knowledge translation tool that leverages summaries generated by an LLM, specifically OpenAI’s GPT-4 Turbo Model, to distill and convey key information from CPGs.
Currently, the prototype, called Guideline Assistant for Evidence based Physiotherapy (GAEP), incorporates two exemplary German national guidelines: one for nonspecific low back pain and another for chronic obstructive pulmonary disease (COPD) (Bundesärztekammer BÄK, Kassenärztliche Bundesvereinigung KBV, AWMF, 2017, 2021). On the homepage, users first select a clinical CPG and enter a query, either as keywords or as a full sentence. The LLM then reformulates the input into a complete question, automatically correcting any spelling mistakes, and generates a concise summary of the key information based on the guideline’s recommendations. Users can also view and sort the original recommendation statements underlying the summary, together with the corresponding source references cited by the guideline authors. To submit a new query, users must return to the homepage.
Figure 1 illustrates the user interface, and a demo version can be accessed.
Upon first use, a tutorial guides users through the tool’s functions, explaining the function of each interface element as it appears. Navigating through the ten pages can be completed in approximately three minutes. The code for the application is available as open source.
Before advancing the development of this LLM-based tool, evaluating its usability and exploring PTs’ perspectives on its potential integration into clinical practice represent key areas of inquiry. While AI applications for knowledge translation are receiving growing attention, little is known about how PTs perceive LLM-powered tools for supporting guideline adherence. To address this gap, the present focus group study was conducted to assess practical feasibility and identify key factors influencing adoption of such a tool. As user experiences and utilization perspectives may differ between settings, a focused analysis was conducted to examine potential differences related to healthcare settings.
The study aimed to investigate PTs’ experiences using the digital tool, explore how they envision its utilization in their clinical context, and analyze perception differences between outpatient and inpatient PT settings. The results of this focus group study provide insights into how an LLM-powered knowledge translation tool could be utilized in everyday practice, informing further development and refinement of the tool to better meet the needs of PTs across clinical settings.
Methods.
Research design
An interpretivist paradigm was employed, using a qualitative design to capture the individual perspectives of the participants. This approach aims to generate a practical understanding of participants’ experiences and the meanings they ascribe to them in context, achieved by conducting interviews and meticulously analyzing the resulting transcripts. Focus group interviews were preferred over individual interviews because they allow multiple participants to be interviewed simultaneously and facilitate social and content-related exchanges. For quality assurance, reporting conformed with the Consolidated Criteria for Reporting Qualitative Research (COREQ) guideline and the Standards for Reporting Qualitative Research (SRQR).
This focus group study was integrated into the final feedback loop of the development process for the In-EbP project, which was registered on the Open Science Framework, and obtained ethical approval prior to data collection. The funders played no role in the design, conduct, or reporting of this study.
Setting and sampling
To obtain insights from both outpatient and inpatient PTs, participants from PT practices and participants from a hospital setting were sought to take part. Therefore, all employed PTs from the project partner PT practices, and the project partner university hospital were eligible for inclusion. The partner practices are a network of eight outpatient clinics that treat a wide range of patients, including orthopedic, neurological, surgical, and sports medicine cases, and employ approximately 70 PTs and related health professionals. The partner university hospital hosts a specialized PT department that provides inpatient and outpatient care for patients with orthopedic, neurological, surgical, and critical care conditions, and employs approximately 46 PTs and lymph therapists.
Purposive sampling was employed, as data saturation is typically estimated to be achieved with four to six focus groups, with three groups being sufficient to identify the most prevalent themes within the data set. Therefore, three focus groups were aimed for in each setting.
Data collection
In April 2024, focus groups were conducted by NLR and DR during working hours of the participating PTs. NLR (male, 29) and DR (female, 34) each had around five years of experience in qualitative research and had previously published on evidence-based practice.
Although not involved in the software development process led by DZ, both were technologically proficient and experienced in working with web-based applications. Both practice as PTs in outpatient settings and had prior familiarity with the participants through their roles as investigators in earlier feedback loops of the In-EbP project. Only the researchers and participants were present. To avoid disturbances, the focus groups were conducted in a separate room at the respective setting. Focus groups began with a 15-minute workshop to test the prototype using a clinical case vignette from a previous study. In these workshops, participants engaged with the case and were asked to answer specific questions related to it (see Supplements). A 30-minute group interview followed, using a pre-tested interview guide (Table 1), developed by DR, NLR, and DZ.
NLR facilitated the workshops and focus groups, while DR was responsible for note-taking and observing nonverbal communication. The interviews were recorded using audio and video equipment and were subsequently transcribed using the transcription software NoScribe (Version 0.4.4). DR pseudonymized and adapted the data to ensure the anonymity of the participants while adding notes on non-verbal communication to the transcripts. The transcripts could not be returned to participants for comment or correction due to feasibility reasons. Before data collection, the research team documented their pre-assumptions, which were set aside and later reflected upon during data analysis.
Data analysis
Following a qualitative content analysis approach as outlined by Kuckartz and Rädiker (2020), transcripts were analyzed using MAXQDA (Release 22.8.0) following both deductive and inductive procedures. In the first step, text passages were deductively assigned to two main categories derived from the research question, focusing on PTs’ perceived experiences and utilization perspectives of the digital tool. In the second step, subcategories were developed inductively from the transcripts.
To ensure rigor, two authors (DR and NLR) independently coded one interview to generate initial coding systems. These were compared, discussed, and merged into a joint draft coding system. To assess coding consistency and refine the categories, both authors then independently coded a second transcript using this draft. Interrater reliability was calculated in MAXQDA using the percentage agreement function, which compares whether text segments were assigned to the same categories by both coders. Discrepancies were discussed and resolved, after which DR coded the remaining interviews, iteratively updating and refining the coding system as needed. The revised subcategories were continuously cross-checked against the transcripts to ensure alignment with the data.
Trustworthiness of the analysis was ensured through multiple strategies. Peer debriefing with BV and DZ supported critical reflection on the coding system and interpretations, clarifying the organization and labeling of categories. Furthermore, a member check was conducted in an online meeting with a subset of focus group participants to validate the findings. Triangulation of data sources, including multiple focus groups and interview notes, combined with iterative coding and reflexive note-taking, provided diverse perspectives on the data, documented analytic decisions and researcher reflections, and helped identify patterns that might otherwise have been overlooked.
As the focus groups were conducted in German, DR translated the cited sequences for publication with the assistance of translation software (OpenAI, Version GPT-4o-mini).
Results.
Participants
Thirteen PTs from an inpatient setting and ten from an outpatient setting in a large German city (> 1 million inhabitants) consented to participate. Participation was facilitated and permitted by their employers, and involvement in the project was voluntary.
Three focus groups in the inpatient setting and one in the outpatient setting were conducted in person, while two outpatient focus groups were held via the online platform Zoom (Version: 6.0.11). The detailed composition of each focus group is illustrated in Figure 2.
As shown in Table 2, participants in both settings were predominantly female, with greater variation in age and work experience observed in the inpatient setting. All participants were professionally trained as PTs, with a subset holding an academic degree.
Coding system
Before reaching agreement, the intercoder reliability of the segments coded into the draft coding system was 48.72%. This primarily reflected ambiguities in the draft coding system rather than substantive disagreements. All differences were resolved through discussion, which supported the refinement and clarification of category definitions. Peer debriefing did not lead to major changes in the overall system but contributed to clarifying the naming and organization of the categories. The member check involved nine participants, representing both inpatient and outpatient settings. While no substantial modifications to the findings were suggested, the feedback facilitated a clearer refinement and clustering of the results, thereby enhancing their clarity and overall trustworthiness.
Data presented as n (%) for each corresponding group (outpatient, inpatient, overall).
As shown in Table 3, the final coding system comprised two deductive categories aligned with the first and second study objectives. The first category was divided into two subcategories, and the second into three. Each subcategory was further broken down into statement codes to facilitate reporting. In the draft version, categories had also been subdivided by setting. However, as results did not differ substantially across settings, the final coding system integrated both settings. To address the third objective, differences between settings were reported only when they were clearly evident within the subcategories. The complete codebook is provided in the Supplements.
Main category: user experience
To address objective one of the research question, the category User Experience focused on the participants’ experiences when engaging with the tool. It was divided into three subcategories: Content Quality, Technical Performance, and Ease of Use.
Regarding the subcategory Content Quality, most participants found the interface design generally appealing, noting the color coding, filters, sorting functions, and pictograms. However, some participants mentioned that the amount of text displayed could be overwhelming for mobile phone use suggesting bolding keywords to improve readability.
It’s good that there is a summary highlighted in color which provides me with the most important information relatively quickly. (Outpatient 3)
Most PTs described the summaries as informative, concise, and accurate, offering quick responses to their queries. While some found the summary details sufficient, others relied on the detailed recommendations for deeper understanding. Though helpful, the detailed texts were seen as time-consuming and used mainly when exploring a topic in more depth.
It was great because you had the precise answer straight away (Outpatient 1)
The AI answer is already very exhaustive. (Outpatient 3)
It’s good that there’s so much more detail underneath. At the top you have this short answer and then you can look more precisely at the bottom. (Inpatient 3)
The content was considered reliable due to the transparency of information sources, the recognizable structure of the underlying CPGs, and the explicit citation of the sources on which the recommendations were based. However, participants felt the content of the CPGs lacked specificity regarding PT. In three focus groups, participants discussed their expectations toward CPGs and emphasized the importance of providing a structured framework allowing for individual application.
Knowing that it comes from a reputable source. That you don’t have to search forever. (Inpatient 3)
It leaves us enough freedom on the one hand but gives us a framework within which we can act. (Outpatient 3)
Compared to alternative sources of information, such as CPGs in PDF format or basic web searches, GAEP was regarded as a more efficient resource. Participants appreciated the access to reliable, scientifically grounded information, contrasting with the time-consuming process of sifting through numerous pages or relying on potentially less credible sources.
It’s better than just typing something into Google somewhere and picking something on the first page because it was clicked on the most. It’s much more scientific, also simply better information. (Outpatient 2)
But when you search on your own and work your way through, I don’t know how many pages of this guideline, it’s incredibly hard to find an answer. This way, you have it all at one click. (Inpatient 1)
In terms of Technical Performance, participants mostly valued the functionalities and the responsiveness of the tool, highlighting the structured layout, the automatic spelling correction, and the keyword search functionality. However, across all focus groups, participants criticized the long response time for the answers. They also expressed concerns about the navigation between different guidelines and aspects of the search bar’s functionality.
I had the feeling that it was organized in a simple and structured way, with the search bar on the top. (Inpatient 2)
The waiting time is still somewhat lengthy to provide a prompt response. (Outpatient 2)
Several recommendations were proposed for improvements, including refining the search functionality by adding options to switch, pause, or terminate the ongoing search, as well as optimizing the guideline selection process by providing brief overviews of clinical cases or categorizing guidelines according to specialties. Additionally, participants suggested adding voice input, incorporating direct links to additional sources, including scientific references and questionnaires, and enabling offline access to GAEP.
It would probably be nice if you could click on a questionnaire, or if you’re specifically looking for tests or exercises, to click on the test to see how it works and don’t have to look it up again. (Outpatient 1)
It could become confusing at some point if there are 30 guidelines. Then some kind of categorisation might be useful, like saying “internal medicine,” and then you have your 10 or so. (Inpatient 3)
Regarding Ease of Use as in comprehensibility and general user-friendliness, most participants indicated that the fundamental functions were readily comprehensible upon first use. However, less technologically proficient participants expressed anxiety when utilizing GAEP.
It was self-explanatory in terms of functions and what to do. I think everyone can still find the search bar. (Outpatient 2)
I’ve generally had a few starting problems where I thought, oh, I’ve got it all wrong. I’m not at all computer-savvy, not even in my private life. And I do have a certain fear of technology. (Inpatient 2)
While less technologically proficient PTs initially struggled to understand the tool, they expected that their uncertainties would decrease with continued use. Similarly, users with higher digital literacy anticipated that repeated interactions with the tool would allow them to explore its various functionalities more effectively. Consequently, repeated use was viewed to enhance the usability of the tool.
Every learning phase simply takes time. But once you’re in there, you get faster and it’s just more fun. (Inpatient 2)
I think it’s just a matter of getting used to it. If you do it often and just use it, then you’ll get used to it quickly. (Outpatient 2)
Opinions on the tutorial varied. Some participants found it overly detailed and too lengthy for understanding the basic functions, while others viewed it as a valuable resource for exploring the tool’s full capabilities. A third group, however, found the tutorial confusing and overwhelming.
The tutorial was short and concise and very easy to use. I really appreciated its user-friendliness. (Outpatient 1)
It took almost too long to click through the tutorial. (Outpatient 2)
It was a bit irritating. Something pops up there. Am I allowed to look up there? (Inpatient 2)
Main category: utilization perspectives
To address objective two of the research question, the category Utilization Perspectives captured statements on how participants envisioned using the tool in their everyday practice. The analysis revealed two key aspects: Determinants of Utilization, referring to factors that may hinder or facilitate its implementation, and Purpose of Utilization, reflecting the contexts in which participants anticipated applying the tool in practice.
The main Determinants of Utilization of the digital tool were found to include the availability of a technical device with internet access and the available time to engage with it. Participants stressed the need for quick results and compatibility with a wide range of devices.
It’s all a question of time. The quicker, the better for us. (Outpatient 2)
It’s best to have as many options as possible, so you can use it on a computer, on a tablet, but also maybe download it onto your phone. (Inpatient 1)
Workplace digitization was another key aspect. Outpatient settings reported better access to devices like tablets, while inpatient settings faced more limited access. However, internet connectivity was stated to be unreliable in both environments.
Everything is being digitized on our ward. There will be more computers available, giving us more opportunities to simply use a computer that’s currently free. (Inpatient 1)
In the next two years nothing will be digitized on our ward. (Inpatient 1)
A problem is that sometimes the internet fails. For example, when we move from cabin 10 to the training room, the internet connection drops. And then it takes time to get it back up. (Outpatient 2)
Inpatient PTs reportedly had opportunities to use the tool during breaks before, between, and after treatments. Conversely, outpatient PTs claimed stricter time constraints, leading to envisioning utilization during periods of inactivity due to treatment cancellations or during the treatment itself. Outpatient PTs commonly used tablets, but in the inpatient setting, this was less feasible due to hygiene concerns. In both settings, the tool would need to be fast enough to prevent any significant loss of therapy time or disruption to the treatment process.
When a patient spontaneously cancels, you can click in the guidelines. (Outpatient 1)
Sometimes it is possible to look at the tablet during treatment to have a quick look. (Outpatient 2)
Outpatient PTs also emphasized that trust between PTs and patients would be essential before using the tool during sessions to maintain the therapeutic alliance and envisioned utilization particularly when patients can perform exercises independently. In contrast, inpatient PTs found it challenging to use the tool with disoriented or immobile patients, as they required constant attention.
90% of my patients aren’t orientated. That’s problematic. I also find it difficult to imagine it in a clinic setting because of the hygiene aspects. (Inpatient 3)
The Purposes of Utilization varied. Participants intended to use GAEP for personal purposes, the patient-therapist interaction, as well as both intra- and interprofessional collaboration. While personal use and patient-therapist interactions overlapped across settings, intra- and interprofessional collaboration differed considerably by context.
On a personal level, participants aimed to utilize GAEP to expand their skills, and acquire up-to-date knowledge, particularly early in their careers. They also envisioned using the tool to address knowledge gaps related to complex medical conditions, cases from other specialties, or to resolve uncertainties. They also saw value in using the tool for therapy preparation and planning, such as supporting clinical reasoning when current approaches prove ineffective. In patient-therapist interactions, GAEP was seen as a valuable tool for patient education, particularly in answering specific patient questions and providing evidence-based explanations to reinforce therapeutic recommendations and enhance the therapist’s credibility.
Maybe when I want something new. (Outpatient 2)
That would definitely be a good help if you are stuck. (Outpatient 2)
I actually see it more as a patient education tool. (Outpatient 1)
At the intra-professional level, outpatient participants viewed GAEP as a tool for quality assurance, standardizing processes, and supporting guideline competencies within the team. They noted its usefulness for onboarding, professional training, and resolving disagreements. Whereas inpatient PTs didn’t comment on its application for intra-professional collaboration, they commented on its use at an interprofessional level, serving as a basis for clarifying responsibilities and understanding other professions’ actions, especially when physician instructions were unclear. Outpatient PTs, however, saw limited utility at the interprofessional level, although they noted it could be useful in improving collaboration with physicians by providing evidence to support the accuracy and appropriateness of prescribed therapies.
It’s also a good way of ensuring quality in PT in general because we don’t have many tools for quality assurance so far. I think this is a simple, quick tool for simply maintaining the internal standard. (Outpatient 1)
In the best-case scenario, you have all the information from the physicians. Or work so closely together. (Inpatient 1)
There is still a need for communication with physicians, even if it’s just to make changes to the prescription. (Outpatient 1)
Differences between settings
Differences between settings were examined during the coding process for each subcategory. The following presents the synthesized results, organized according to the two main categories of the coding system, thus allowing for a clear comparison of participants’ User Experiences and Utilization Perspectives of GAEP across inpatient and outpatient contexts.
Participants’ User Experience with GAEP were largely similar in terms of content quality, technical performance, and ease of use. No major differences emerged between inpatient and outpatient settings regarding the perceived reliability, structure, or comprehensibility of the tool. Variations were primarily related to participants’ digital literacy rather than the clinical setting.
Differences between settings became more apparent when considering the Utilization Perspectives of GAEP in practice. Outpatient PTs reported stricter time constraints and emphasized the need to integrate the tool during short periods of inactivity or with patients capable of independent exercises. They also highlighted its use for intraprofessional purposes, such as quality assurance, standardization of practices, onboarding, and team training. In contrast, inpatient PTs stated having more opportunities to use the tool during breaks and focused on interprofessional collaboration, for instance, clarifying roles or understanding instructions from physicians. Device availability and technical infrastructure also differed.
Outpatient settings generally provided easier access to tablets, while inpatient settings were limited by hygiene requirements, although internet connectivity issues were reported in both contexts. Across both settings, GAEP was considered useful for personal knowledge acquisition and patient education, but the specific applications within team and interprofessional contexts varied according to the setting.
Discussion.
The main findings indicate a positive overall user experience, with GAEP efficiently responding to questions arising in daily PT practice. Participants found the interface appealing and the summaries concise and informative, providing quick access to reliable, evidence-based content. The tool’s structured layout, search functionality, and intuitive basic functions were valued, and repeated use was expected to further improve usability, with the tutorial supporting exploration of its full capabilities. In terms of utilization, GAEP was reported to mainly depend on available time and access to internet-enabled devices. Participants anticipated using the tool for various purposes, including personal use, patient-therapist interaction, and both intra- and interprofessional contexts.
No considerable differences in user experience were noted between settings, but experiences seemed to depend on digital literacy. Differences between settings were most apparent in utilization perspectives, with outpatient PTs, constrained by time, emphasizing use during brief inactive periods or with independent patients and for intraprofessional purposes, whereas inpatient PTs had more opportunities to use the tool during breaks and focused on interprofessional collaboration. Device availability and technical infrastructure also differed between settings, yet GAEP was valued in both contexts for personal knowledge acquisition and patient education.
The findings provide insights into the perception and utilization perspectives of an LLM-based web application in the context of German PT, highlighting the potential of such tools in improving access to and adoption of CPGs. These findings align with the growing body of research on the application of artificial intelligence in healthcare, particularly in knowledge translation and clinical decision support tools. For instance, AI-enabled clinical decision support systems have been shown to influence healthcare adoption behaviors through factors such as performance expectancy, effort expectancy, social influence, facilitating conditions, and trust, while at the same time enhancing decision-making, diagnostic accuracy, and workflow efficiency.
Participants noted the overall advantages of GAEP compared to PDF versions of CPGs, indicating that barriers to guideline adherence, such as text length, usability, and accessibility to CPGs, could potentially be addressed by an LLM-based tool. Although quantitative comparisons remain necessary, such tools appear to offer distinct advantages. Yet this is still acceptable for a prototype, GAEP was criticized for its slow performance and the insufficient number of integrated CPGs, currently limiting its practicality for clinical use. With further development, especially optimization of speed and expansion of integrated content, the findings indicate promising potential for a highly practical tool for supporting guideline-based decision support in fast-paced healthcare settings.
While many PTs have positive expectations toward new technologies such as LLM-based tools, others face challenges with technological advancements. In this focus group study, participants with lower technological proficiency reported more negative experiences underscoring the importance of considering healthcare professionals’ digital literacy when implementing such tools. This finding is consistent with research on the adoption of digital health intervention tools, which identifies digital literacy as a critical factor for successful implementation. Frameworks on eHealth Literacy suggest that the effective use of AI-driven tools requires not only technical skills but also the ability to critically evaluate and apply digital information.
Therefore, structured training programmes focusing on AI literacy and CPG interpretation should be developed to enhance PTs’ confidence and competence in using LLM-powered tools in clinical practice.
Some participants criticized the lack of PT-specific content integrated into the tool, which they felt limits its current utility. This finding suggests a need among German PTs for more tailored CPGs. As previously outlined, CPGs in Germany are typically not specifically designed for PT. While this situation may evolve in the future due to increasingly interdisciplinary guideline development groups, it still presents a challenge. PTs in Germany tend to prefer tools specifically designed by and for their profession. Therefore, incorporating international PT-specific CPGs into GAEP could provide more detailed information pertinent to the field. Yet, CPGs are designed to align with the specific healthcare system and setting for which they are developed (Institute of Medicine (US), 2011; 1990).
As a result, including international CPGs could compromise the tool’s accuracy within the local geographic context, given the unique aspects of healthcare systems, regulations, and clinical pathways. On the other hand, the tool’s development could be expanded to other healthcare professions offering an effective strategy for enhancing guideline adherence, since barriers to guideline adherence are similar across various fields.
The reported determinants of utilization, such as time constraints and access to internet-enabled devices, align with findings from studies on clinical decision support tools, telehealth interventions, and the adoption of evidence-based practice in PT. Workplace digitization, identified as a key determinant, should be assessed before implementing digital tools like GAEP. This supports the recommendation that strategies be specifically tailored to the context in which innovations are to be implemented. Therefore, to facilitate GAEP’s integration into PT workflows, it seems essential to consider its adaptability to different practice settings. For instance, in the outpatient setting, a mobile-friendly version with voice-activated functionality could improve accessibility during patient interactions.
In inpatient settings on the other hand, integrating GAEP with electronic health records (EHRs) could streamline decision-making by embedding real-time, evidence-based recommendations within existing documentation systems. Future research should assess the feasibility of these adaptations and their impact on workflow efficiency and patient outcomes.
However, evidence shows that dissemination alone rarely ensures adoption or integration of recommendations into clinical practice (Damschroder, Reardon,
Widerquist, and Lowery, 2022; Powell et al., 2015). Thus, simply making such a tool available is unlikely to directly change behavior or improve guideline adherence. Successful uptake typically requires active implementation support, including trained specialists who assess the local context, identify barriers, and deliver tailored strategies, as recommended by established implementation frameworks.
Other challenges arise with LLM-powered tools. Interpreting CPGs through an automated system carries a risk of misinterpretation when taken out of context. CPGs are developed as a comprehensive framework to illustrate a clinical picture and cannot directly function as clinical decision support systems (Berner, 2016; Greenfield, 2017; Institute of Medicine US Committee to Advise the Public Health Service on Clinical Practice Guidelines, 1990). Another concern is that even though the tool was intended to extract information solely from the underlying CPGs, there remains a potential risk of generating inaccurate or “hallucinated” responses, a concern often associated with LLM-based tools.
Beyond hallucination, which may diminish over time as models advance and users become more skilled in interacting with them, broader ethical issues must be considered when implementing AI in clinical practice. This includes evaluating the tools’ accuracy, privacy issues, informed consent, and responsible use. However, this issue was not raised by participants in our study. They generally expressed positive perceptions of GAEP’s content quality and reliability, showing little skepticism toward the responses provided by the tool. This absence of concern may reflect the controlled study context, yet ethical safeguards remain crucial for practice. To mitigate these risks, LLM-powered tools would have to undergo validation. Such validation should not only focus on technical accuracy but also integrate ethical standards for responsible deployment.
This process is complicated by the rapid evolution of technologies, as tools may become outdated by the time they are validated. Furthermore, regarding validation, establishing a gold standard for comparison, such as human expertise, is debated in the literature.
Participants envisioned using GAEP not only to access CPGs but also to support clinical reasoning in complex cases. This introduces risks, particularly when users apply the tool in areas beyond their expertise, such as utilizing assessments from different specialties. This raises an important question: Can LLM-based tools genuinely expand users’ competencies, or do they risk fostering over-reliance on technology, potentially diminishing critical thinking and clinical judgment? Similar concerns have been highlighted in broader discussions about artificial intelligence in healthcare, where the balance between augmenting human decision-making and promoting dependency on machines remains unclear.
Moreover, although not extensively discussed in the results of this study, the implications of using technical devices in healthcare warrant further exploration. Participants briefly touched on concerns regarding the impact of the tool on the patient-therapist relationship. The use of tools like GAEP may alter the dynamics of this relationship, potentially affecting trust, transparency, and the balance between human judgment and technological assistance. Future research should address these concerns, particularly how to ensure that such technologies enhance rather than undermine the therapeutic alliance.
Furthermore, outpatient PTs envisioned using GAEP for quality assurance and standardization of treatments. Thereby, the tool could help reduce practice variations and improve adherence to evidence-based practice, a goal that aligns with broader efforts to enhance healthcare quality. However, research is needed to determine whether LLM-based tools like GAEP can consistently improve PTs’ guideline adherence and reduce discrepancies in care delivery across settings. Additionally, exploring the role of these tools in promoting standardized care without compromising individualized patient treatment represents another direction for future research.
Limitations.
Only a few interviews were double-coded, which may introduce bias into the results due to the personal perspectives of the primary coder (DR). Additionally, the interpretation of the researchers and the statements of the participants might have been biased, as the researchers were involved in the development of the tool and were familiar to the participants from earlier feedback loops. Furthermore, no theoretical framework or model was used to analyze the findings, and the analysis did not account for the level of agreement across participants or focus groups. While some statements were endorsed by a few PTs, others had broader support, leading to variation in their weight. Moreover, in some focus groups, the small number of participants hindered the natural flow of conversation, resulting in a format similar to individual interviews.
While this constraint affected group dynamics, directly opting for individual interviews rather than focus groups could have been more appropriate, considering the study’s focus on exploring individual behavior change.
Conclusion.
In this focus group study, the experiences of PTs using GAEP, an LLM-based knowledge translation tool, and their expectations for its utilization, were analyzed. The findings indicate that LLM-powered tools hold promise for enhancing guideline adherence among German PTs, as participants across both settings identified a wide range of potential purposes of utilization. However, GAEP is not yet fully usable, requiring further development particularly to improve response speed and incorporate additional CPGs. Implementation strategies of such tools should be tailored to the specific clinical setting and account for its digitization status as well as the users’ digital literacy. Expanding the tool to international CPGs and other healthcare contexts could be explored.
Future research should focus on quantitatively validating these qualitative insights, validating the reliability of LLM-generated summaries, and assessing whether LLM-based tools can effectively improve guideline adherence. Ethical considerations, such as the impact on clinical reasoning, competencies, and the therapeutic alliance, remain critical in the development and integration of such technologies.
Acknowledgments.
We would like to thank the participants for their insights, our project partners for ensuring the availability of the participants, IFAF Berlin and the ‘Senate Department for Science, Health, Care and Equality, Berlin for funding the project, as well as Dr. Karen Otte, Prof. Christiane Stock, Prof. Hürrem Tezcan Güntekin, and Dr. Verena Struckmann for their methodological assistance. Furthermore, we would like to thank Leonard Rosen for proofreading the manuscript.
CRediT: Diane Rosen: Conceptualization, Data curation, Formal analysis, Investigation, Methodology, Validation, Visualization, Writing – original draft, Writing – review & editing; Dorian Zwanzig: Conceptualization, Data curation, Methodology, Resources, Software, Writing – review & editing; Barbara Vogel: Conceptualization, Funding acquisition, Resources, Supervision, Validation, Writing – review & editing; Michael Erhart: Conceptualization, Funding acquisition, Project administration, Supervision, Writing – review & editing; Nils Lennart Reiter: Conceptualization, Formal analysis, Funding acquisition, Investigation, Methodology, Project administration, Validation, Writing – review & editing.
Author contributions
Disclosure statement
The authors reported no conflicts of interest according to the ICMJE Form for Disclosure of Potential Conflicts of Interest. However, they received funding for the development of the prototype, which may be considered an intellectual conflict of interest.
Funding
This work was supported by the Institute for Applied Sciences Berlin (IFAF Berlin) and the Senate Department for Science, Health, Care and Equality, Berlin.
AI statement
The translation of citations from the transcripts, as well as the grammatical and spelling optimization of the article text, were carried out with the assistance of AI tools (DeepL and OpenAI). Adoptions were made as needed, and the authors take full responsibility for the content.
Data availability statement
The transcripts that support the findings of this study are not publicly available but may be available from the corresponding author upon reasonable request.
Ethics approval
The project was approved by the Ethics Committee of the Alice Salomon University of Applied Sciences Berlin (Approval Number: 03–2023/54).