1 More Paper.
Full Reading01:21:31

Generative artificial intelligence in creative contexts: a systematic review and future research agenda

1 More Paper · Full Reading

Full Reading podcast cover
Listen to the Full Reading

About this paper

A full audio edition of this paper.

Authors: R. Heigl

Publication date: 2026

Read the paper: https://doi.org/10.1007/s11301-025-00494-9

Source license: Creative Commons Attribution 4.0 International — https://creativecommons.org/licenses/by/4.0/

The authors and publisher do not sponsor or endorse this recording.

Brief episode

Transcript

You’re listening to “Generative artificial intelligence in creative contexts: a systematic review and future research agenda,” by R. Heigl. Published in 2026.

Abstract.

Generative artificial intelligence (GenAI) has recently attracted attention from litera-ture and organisations, especially due to advances in machine learning techniques. However, research on GenAI in creative contexts remains in its early stages, with few attempts made to assess the current body of research or synthesise the exist-ing knowledge in this area. To address this gap, this paper employs a systematic literature review of 64 studies to identify methods, research trends and key thematic insights shaping the current understanding of GenAI in creative contexts. The find-ings of this systematic literature review emphasise the rapid development of research on GenAI in creative contexts. The analysis highlights key factors influencing the adoption and impact of GenAI in creative processes, as well as the implications for creative outcomes and industry practices.

From this analysis, several potential direc-tions for future research emerge, including the long-term effects of GenAI on crea-tive processes, socio-economic implications for creative industries, and frameworks for ethical use, and perception of GenAI-generated content.

Over the past two years, generative artificial intelligence (GenAI) has transformed creative industries and introduced a new era in which artificial intelligence (AI) sys-tems collaborate with humans to produce innovative and revolutionary results. From music compositions and fine arts

, to literature and design, GenAI has moved the boundaries of creativity. The Museum of Modern Art’s (MoMA) acquisition of Refik Anadol’s Unsupervised series marks the first time a museum has acquired an GenAI-generated artwork for its permanent collection. Initiatives to promote GenAI-generated art, such as Refik Anadol’s Unsupervised series, underline the growing importance and acceptance of GenAI-generated artworks in the art world. With the increasing development of GenAI and the quality of its artistic output, it can produce art that is indistinguishable from human-made works. For example, GenAI-generated art-works produced by systems like DALL-E and Midjourney can be highly photoreal-istic, capturing subtle details and nuances once thought exclusive for human artists. Consequently, even professionals are having a hard time detecting the true source of artworks.

In line with this assertion, research started to point out that humans cannot distinguish between art generated by GenAI and humans. Nascent research argue that the emergence of GenAI is revolutionising the meaning of creativity, challenging the notion that creativity is a uniquely human ability. The introduction of GenAI to creative industries was the topic of several lit-erature reviews, each analysing different aspects of this growing field. For instance, Carnovalini and Rodà (2020) reviewed computational creativity, explicitly focus-ing on music generation. Anantrasirichai and Bull (2022) provided insights into the adoption of GenAI in creative industries from a technical perspective. Oksanen et al. (2023) reviewed AI tools and their utilisation in fine art production. Another study by Ameen et al. (2022) explored theories of AI for creativity in marketing.

Despite these individual reviews, there has yet to be a systematic literature review (SLR) of current research on creativity and GenAI that presents a comprehensive over-view and establishes a guideline for future research directions. Further, this review attempts to bring together the perspectives of artists and how GenAI influences their creative process with the perception response of recipients who consume content created by GenAI, thus bridging the gap between content creation and consumption in the context of creativity. Finally, based on the thematic analysis, this review offers a conceptual framework that explores both the perception of GenAI and its associ-ated factors in creative contexts, as well as the collaborative process between artists and GenAI systems. In particular, this review is guided by three research questions:

1. What themes around GenAI in creative contexts have been identified and exam-.

ined to date by researchers? a. Can individuals distinguish between human and GenAI-generated art? b. How do individuals evaluate GenAI content in creative contexts? c. How do artists work with GenAI in creative contexts?

2. What measurements are used to evaluate content generated by GenAI?

3. Which issues need to be addressed in future research?

To answer these questions, this study employs a SLR with analysis of 64 articles with publication dates from 2015 to 2024. This study offers a comprehensive review of GenAI and creativity through literature. The main contribution of this study is to offer insights into the current state of literature, and outlining potential directions for future research. The structure of this review is as follows. First, there will be an introduction to previous reviews related to GenAI and creativity. Second, the con-cepts of GenAI and creativity are elaborated. Next, the review methodology is out-lined. This is followed by a descriptive overview of the findings. Subsequently, the findings of the selected articles are systematised and analysed. Finally, the author proposes avenues for future research and discusses the implications, while acknowl-edging the limitations of the study.

2 Related literature 2.1 Previous reviews on GenAI in creative contexts

While the rapid advancement and accessibility of tools like DALL-E and Midjour-ney have started to change the creative landscape, the literature on this topic remains relatively nascent. Few SLRs have conducted comprehensive citation and content analyses on this topic. In response, this study explores this research gap and identi-fies shortcomings in GenAI research within the Information Systems (IS) domain. While acknowledging four prior reviews, the author also points out their limitations and seeks to address them through a more extensive and sys-tematic review of the literature (see Table 1).

Anantrasirichai and Bull (2022) provide a technical analysis of GenAI’s integra-tion into creative industries, detailing how advancements in machine learning and neural networks are revolutionising workflows in visual effects, animation, and digital content creation. They conclude that GenAI’s value lies predominantly in its ability to enhance, rather than replace, human creativity. Carnovalini and Rodà’s (2020) review offers a comprehensive exploration of computational creativity, with a particular focus on the realm of music generation. Their research not only shows the complex algorithms and AI techniques driving music production but also under-scores the dual role of AI as both collaborator and creator within the music industry.

However, their review was published in 2020 and was therefore unable to capture recent findings on this fast-growing topic such as including the launch of ChatGPT or DALL-E. Neither Carnovalini and Rodà (2020) nor Anantrasirichai and Bull (2022) adopted a SLR methodology, and both omitted a detailed account of the final number of articles included in their analyses. Oksanen et al. (2023) turn their attention to the fine arts, investigating in 44 articles how art-ists are embracing GenAI tools to test artistic boundaries to create innovative works. While this review addresses the production aspects of employing GenAI in crea-tive fields by identifying the tools utilized, it does not consider artists’ perspectives on how their creative processes are influenced by the use of GenAI. In a similar vein, Ameen et al. (2022) researched the application of AI in marketing creativity,

Table 1 Comparison of previous reviews on GenAI in creative contexts constructing a theoretical framework that elucidates how AI-driven innovation can elevate brand storytelling, foster deeper customer engagement, and refine personal-ized marketing strategies. Collectively, these studies illustrate the multifaceted influ-ence and potential applications of GenAI in creative contexts. However, considering the rapidly growing progress and implementation, it is necessary to revisit current literature to cover latest developments and findings.

2.2 Generative artificial intelligence

GenAI can be defined as a technology that (i) leverages deep learning models to (ii) generate human-like content (e.g., images, words) in response to (iii) complex and varied prompts (e.g., languages, instructions, questions). GenAI represents a significant advancement in AI, departing from traditional predictive models. Unlike predictive AI, which focuses on analysing input data to make predictions or decisions, GenAI aims to create new content or data similar to its training data but not identical. Predictive AI models, such as those used in classification or regression tasks, extrapolate from existing data to predict outcomes. These models excel in structured environments in which the goal is to understand patterns or trends within the data. In contrast, GenAI does not just interpret data but actively generates new content, in a creative or inno-vative way.

This generative process involves understanding and replicating the underlying distribution of the training data. The capability of produc-ing novel outputs makes GenAI a powerful tool for various applications, including artistic creation, data augmentation, and problem-solving in novel ways. One can broadly distinguish between three types of GenAI systems: (i) text-to-text, (ii) text-to-image, and (iii) text-to-music. Text-to-text systems accept text input (prompt) to produce a natural language output, e.g., text, code, or even music. Examples include machine translation, summarization, and question-answering systems. While models like ChatGPT learn to predict the next word, producing coherent text, text-to-image like DALL-E systems translate text to visuals. Finally, text-to-music systems like MusicLM translate text to audibles.

Further, ethical concerns have risen around copyright, authenticity, and potential misuse for misinformation. This concern adds another layer of complexity to the ethical considerations surrounding GenAI-generated art, including issues of originality and the value of human creative labour. Also, GenAI output can perpetuate or amplify biases present in training data, raising questions about fairness and representation. Additionally, these systems’ substantial computational requirements and energy consumption present environ-mental challenges. Ultimately, the integration of GenAI into cre-ative contexts represents a technological shift and, more importantly, impacts the relationship between human creativity and artificial intelligence.

2.3 Creativity

Defining and measuring creativity has been a complex and multifaceted endeavour. In 1950, Guilford (1950) highlighted a lack of research to the American Psychologi-cal Association. Since then, increasing studies have been conducted to explore the concept of creativity. While scientists initially believed creativity was solely psychological, its academic importance increased. Over the years, defining and measuring creativity has proven challenging due to its multifaceted nature. Treffinger (1996) extensively reviewed creativity literature, revealing over 100 definitions of the concept. Moreover, researchers and educators often interpret creativity differently, associating it with various cognitive processes, personal attrib-utes, and past experiences.

Additionally, terms such as inno-vation, invention, imagination, talent, giftedness, and intelligence are sometimes used interchangeably with creativity. Table 2 shows exam-ples to showcase the variety of definitions of creativity.

In general, scholars refer to four main perspectives to define creativity: the cogni-tive processes associated with creativity (“process”), the individual characteristics of creative people (“person”), the creative results or outcomes (“product”), and the interplay between the creative individual and their environment (“press”). The establishment of the 4P construct reveals an interplay of individual psychological factors and environ-mental influences for creativity. Creativity is not solely generating a novel product but emerges from the individuals’ intentional processes and environment. This study will refer to creativity as the capacity to generate ideas or solu-tions that are not only novel and original but also meaningful and valuable within their specific context, arising from an interplay of cognitive processes and environ-mental factors.

3 Methodology

In line with the research questions, the author follows a systematic literature review.1 A systematic review reduces errors and biases, enables researchers to identify gaps, to propose additional research, and deepen understand-ing. Additionally, this approach enhances overall quality by employing a transparent and easily reproducible procedure. Following Tranfield et al. (2003) and Crossan and Apaydin (2010) the author proceeded in three steps: data collection, data analysis, and synthesis. In the first step, the author identified the research objective and formed research questions before she started the literature search. The data collection was carried out across multiple databases to cover a wide range of studies from vari-ous journals. The online databases used were Web of

Science, ScienceDirect, and Google Scholar. Web of Science and ScienceDirect were chosen for their academic rigor and focus on relevant disciplines, while Google Scholar complements these by capturing litera-ture through its broad indexing scope, which allows for the inclusion of new or less common perspectives that may not yet be captured in traditional databases. Based on the recommendation of Khosravi et al. (2019) initial searches on ScienceDirect and Google Scholar were conducted to determine relevant keywords. The author constructed the keywords based on the key components of the research questions and related to the identified key attributes of GenAI in creative contexts.

The follow-ing search strings were searched within the studies’ titles, abstracts and keywords in the databases: “Generative artificial intelligence” OR “Generative AI” OR “Genera-tive machine learning” AND “Creativity” OR “Computational creativity” OR “Cre-ative thinking" AND “Art” OR “Music” OR “Fine art” OR “Creative writing” OR “Visual art” OR “Poetry”.

Following previous studies, the author identified inclusion and exclusion criteria to ensure that the chosen studies are rel-evant for the purpose of the systematic review. Only studies that fulfilled the follow-ing criteria were included:

(a) peer-reviewed,

(b) published in English or German language,

(c) published between 2015 and 2024.

The inclusion of English and German studies was determined by the author’s lan-guage proficiency and the predominance of these languages in accessible scholarly literature on the topic. The limited time frame arises from the objective to cover the latest literature on this emerging topic.

Studies that were excluded have.

(a) not been available as full text,

(b) not been available in the English or German language, and

(c) as the focus of this review is on creativity, only studies dealing with creative or artistic topics have been included, and those not relating to GenAI in creative contexts have been omitted. Additionally, studies focusing on purely technical aspects of GenAI without a clear connection to creativity were excluded. The author also omitted reviews, meta-analysis studies, and non-refereed publications.

Nevertheless, the literature search may have unintentionally omitted some rele-vant articles due to database limitations or researcher oversight. Additionally, variations in terminology across disciplines could lead to relevant studies being overlooked, as the SLR outcome depends heavily on the keywords available at the time. However, the author tested the search terms and keywords on a smaller sample to make adaptions to the search strategy if necessary.

First, 1727 studies were retrieved by using the defined search terms in Web of Science and Science Direct. 430 duplicates were sorted out. Second, the author screened the studies’ abstract and title, removing 925 studies that did not match the research questions. Third, the author full-text screened the retrieved 372 stud-ies independently to determine whether the inclusion criteria were met. In this step, 321 studies were removed. Finally, the author performed a backward-for-ward search to find more relevant studies by using Google Scholar and ProQuest. The backward search involved reviewing references from the initially chosen studies to identify additional relevant articles, while the forward search used cita-tions of these studies to find newer research.

For the backward-forward search the author reviewed the references from the chosen studies and included 13 more articles that met the inclusion criteria. In total, 64 articles matched all criteria and were identified as final sample for the reporting (see Appendix for the com-plete list). The coding process involved categorising articles based on bibliomet-ric details and thematic focus, including GenAI tool, art form, and measurement approaches. Due to the absence of an existing suitable coding scheme in this emerging field, a multiple classification framework was developed. This coding framework was designed to be flexible and iterative. Figure 1 illustrates the selection process.

4 Results 4.1 Descriptive details on the selected articles

The analysis shows a trend in the area of research on GenAI in creative contexts. As shown in Fig. 2, scientific attention to this topic has increased significantly in recent times. The temporal distribution of publications on GenAI shows a clear upward tra-jectory. The analysis shows that half of the 64 articles in the dataset were published within the last year, which underlines the fast pace of research in this area. Consid-ering the revolutionary impact of GenAI, exemplified by the release of DALL-E in January 2021 and ChatGPT in November 2022, it is not surprising that the analysis shows a concentration of publications within the last three years. The most articles were found for 2023 with n = 30 articles, followed by 20242 with the second most articles n = 13, and 2022 with the third most arti-cles n = 9.

The articles were categorised into research themes based on their focus to address the research questions. The analysis revealed two primary themes: Art-ist perspective on working with GenAI and perception of GenAI-generated content. The “Artist perspective” theme, comprising 30 articles, explored how artists work together with GenAI and how it influences their creative process. The largest cat-egory, “perception of GenAI-generated content”, included 34 articles examining various aspects of how users perceive and interact with GenAI-generated content. Following the approach of recent systematic reviews in the field, this study emphasizes key findings rather than providing an in-depth analysis of each individual article. The next sections will aim to address the research questions in detail, examining key themes, applied measurements, and insights drawn from the reviewed articles.

Further, the findings of this research are synthesised into a comprehensive, multi-dimensional framework for GenAI from both the artists’ and the recipients’ perspectives (Fig. 5).

4.2 Can individuals distinguish between human and GenAI‐generated art?

In recent years, research has started to explore the ability of humans to distinguish between GenAI-generated art and human-made art. Several researchers have applied the Turing Test to investigate this (see e.g., Badura et al. 2022; Gun-ser et al. 2022; Horton et al. 2023; Köbis and Mossink 2021; Mazzone and Elgam-mal 2019). The application of the Turing Test across various forms of artwork, such as music, poetry, and visual art, has yielded mixed results in determining whether humans can reliably distinguish between human-made and GenAI-generated crea-tions. In a creativity-focused Turing Test, researchers assess the ability of GenAI systems to generate creative content that is indistinguishable from human-made art. This process typically involves four steps. First, GenAI systems are tasked with pro-ducing creative outputs such as artwork, music, or poetry.

Simultaneously, human artists create similar works. Next, a panel of human judges is presented with a mixed collection of both GenAI-generated and human-made artworks, without knowing the origin of each work. The judges are then asked to attempt to distinguish between the GenAI and human-made pieces, often providing explanations for their choices. If the human judges are unable to detect the true origin of the artworks, the GenAI is considered to have “passed” the test. This outcome would suggest that the AI has successfully produced creative content that is, at least on a surface level, com-parable to human creativity in its ability to engage and convince human observers. In visual arts such as paintings, the ability of individuals to detect the true origin of artworks—whether human-made or GenAI-generated— is influenced by the nature of the artwork itself.

Chamberlain et al. (2018) found that type of artwork matters for an individuals’ ability to detect the true origin. Individuals showed a bias towards images labeled as human-made, especially for representational images compared to abstract images. Corroborating this finding, Gangadharbatla (2022) suggests a cog-nitive heuristic where abstract art is linked to machines, while representational art is associated with human creators. This cognitive heuristic underscores the role of stylistic elements in shaping viewers’ assumptions about the origin of artwork. Inter-estingly, individuals were significantly more accurate in identifying human-made art (64%) compared to computer-generated art (40%). However, Samo and Highhouse (2023) provided empirical evidence that people cannot distinguish between human-and GenAI-generated images.

For poetry, Köbis and Mossink (2021) also leveraged incentivized Turing Tests to explore the extent to which individuals could accurately detect the origin of poems. Only Köbis and Mossink (2021) experimentally varied the selection process for the content to be distinguished from the individuals. Their findings revealed an additional factor influencing human detection capability: the entity responsible for selection. While people could distinguish randomly chosen poems from GenAI-generated ones when selected by humans, they struggle to do so when the poems have been curated by a large language model (LLM). Notably, no other study has included LLMs in the selection of artworks yet. Badura et al. (2022) provided evidence suggesting that individuals were unable to correctly iden-tify the true origin of written text.

However, Badura et al. (2022) acknowledged a limitation in their study design, noting that the brevity of the content presented to participants may have impeded their ability to make accurate judgments on the true origin of the artwork. Additionally, a GenAI system for automatically generating poetry based on images has succeeded in passing the Turing Test. These studies demonstrate the utility of Turing Tests in assessing the capabilities of GenAI systems and the challenges associated with distinguishing between human and machine-generated output. It suggests that our ability to distinguish between human and GenAI-generated art is not an isolated skill, but rather a multi-layered cognitive process that adapts to the unique characteristics of each artistic medium.

In summary, the ability to recognise the origin of creative works seems to depend less on the overarching artistic field than on specific characteristics peculiar to each form of art. The collective results of these studies suggest an increasing conver-gence between GenAI-generated and human-made art and question our ability to distinguish between the two. Humans do have some capacity for distinguishing, but this ability is imperfect and susceptible to various contextual influences. As GenAI advances, this distinction is likely to become increasingly difficult, potentially requiring a fundamental reassessment of our understanding of artistic creation and appreciation.

4.3 How do individuals evaluate GenAI content in creative contexts?

4.3.1 Impact of labeling on art evaluation

This section explores the evaluation of art created by GenAI. A consistent theme emerging from the literature is the influence of labeling on art evaluation. Multi-ple studies have demonstrated that simply labeling a piece as GenAI-generated or human-made significantly affects its perceived value and quality (see e.g., Hong and Curran 2019; Ragot et al. 2020; Köbis and Mossnik 2021). The majority of stud-ies reveal a preference for human-made art. Ragot et al. (2020) and Millet et al. (2023) found that human-made art was consistently perceived as more creative and awe-inspiring. Bel-laiche et al. (2023) corroborated these findings, noting higher ratings for enjoyment, beauty, depth, and value in art labeled as human-made. This bias was particularly strong among individuals who viewed creativity as a uniquely human trait.

Similarly, Gunser et al. (2022) reported that GenAI-generated content was perceived as less well-written, inspiring, fascinating, interesting, and aestheti-cally pleasing than human-written pieces. Interestingly, this negative bias can per-sist regardless of whether participants were informed about the poem’s algorithmic origins, indicating favouritism for human-made content. In contrast, Zhang and Gosline (2023) found that the preference for human-made content persists, primarily due to human favouritism rather than algorithm aversion. In other words, when people know that a human made the content, they tend to rate it as higher quality, but knowing that GenAI played a role in creating content did not significantly lessen their perception of it.

While Demmer et al. (2023) found no significant difference in perceived quality between GenAI- and human-made art, GenAI-generated artworks were rated as less emotional and intentional, leading to lower purchase intentions.

4.3.2 Contextual and individual differences

Further, art co-created with GenAI is seen as more innovative but less authentic and labour-intensive, which diminishes its value in the eyes of the audience, especially in art and culture. However, this nega-tive perception can be mitigated by disclosing the level of human involvement in the creation process. In contrast, Magni et al. (2024) found that individuals perceive creative processes as less effortful when conducted with GenAI rather than by humans, resulting in lower assessments of the creativity of GenAI-generated content. Even when GenAI systems put in more effort, human-made art with less effort is still perceived as more creative, underscoring that the creator’s identity has a more significant impact on perceived creativity than the level of actual effort involved.

However, this effect var-ies by context, with no significant differences in creativity ratings for GenAI-generated content in advertisement posters and name ideas but noticeable differ-ences for paintings. The perceived commercial intent behind artistic creation appears to mitigate the negative impact of GenAI involvement on audience reception. When art is viewed through a commercial lens, the negative effect diminishes. Moura et al. (2023) observed that GenAI-generated content, particularly for intangible products, was per-ceived as higher quality due to the novelty associated with its creation process. Additionally, Gangadharbatla (2022) demonstrated that individuals are more inclined to evaluate abstract artworks favourably when attributed to computers, whereas representational artworks receive less favourable evaluations under the same attribution.

While most studies indicate a bias against GenAI-generated art, Hong and Cur-ran (2019) observed that participants’ evaluations of art were generally unaf-fected by whether they knew the work was GenAI-generated. However, they noted that individuals with strong beliefs about GenAI’s limitations consistently rated GenAI-created pieces lower, underscoring the role of preexisting beliefs in shaping perceptions. While people with strong beliefs about GenAI’s limitations tend to rate GenAI-generated art lower, affirming messages aligned with their beliefs can reduce these negative attitudes. For individual differences, Heigl et al. (2023) found that men tend to trust human-labeled con-tent more, while women show a greater tendency to trust GenAI-generated con-tent, regardless of labeling, highlighting possible gender differences in informa-tion processing.

Lattika et al. (2023) discovered that competence in using GenAI does not correlate with positive attitudes toward GenAI-generated art, suggest-ing that familiarity with GenAI does not always translate into acceptance in artistic contexts. Further, the studies from this sample have largely focused on evaluating GenAI-generated content through controlled experiments carried out in laboratory environments (see e.g.; Köbis and Mossink 2021; Messer 2024). Based from the SLR findings, the author developed a conceptual model (Fig. 3) to articulate key propositions and structure the insights.

4.4 How can artists work with GenAI in a creative setting?

4.4.1 Redefining creative workflows with GenAI

The integration of GenAI into creative workflows has initiated a shift in the design process that can be differentiated from traditional art. Seidel et al. (2018, p. 51) describe this phenomenon as the “triple-loop approach”, where the interaction between humans and autonomous computational tools such as GenAI actively shapes the design outcomes as well as the design process. In contrast to traditional design tools, which typically serve a supportive role, GenAI is the first tool to take an independent role in influencing the creative process.

The triple-loop approach to design involves three key elements. First, the tool itself generates design alternatives based on a set of input parameters and predefined evaluation criteria. Second, this generates a feedback loop which enables the artist to evaluate the alternatives and to modify the tool’s inputs and settings to improve the design process. Within this process, Lyu et al. (2023) and Verheijden and Funk (2023) point out that artists working with GenAI tended to get caught in a loop of excessive refinement and organisation, which can inhibit unconventional thinking and lead to over-structured outcomes. Setting a time constraint can help to break this pattern by preventing endless idea refining. Alterna-tively, the tool can engage in machine learning to autonomously refine its own model and generate better alternatives.

The third loop involves mutual learning, where the human artist gains insight into the tool’s embedded mental models, while the tool learns about the artist’s own thought processes, allowing for alignment between the two. Research revealed that artists experience a positive impact on their own creative process through this approach, particularly for idea generation and inspiration. Artists can input prompts or initial concepts into GenAI systems, which then generate a multitude of variations or interpretations. This can lead to new ideas and to unexpected creative directions that the artist might not have considered otherwise. However, while collaborative ideation with GenAI may enhance originality and nov-elty, it often failed to align with individual aesthetic tastes, lacks adaptive creativity and depends on the artist’s skill set.

Kim (2024) corroborates this finding by revealing that artists perceive that they cannot put sociocultural depth and individu-ality into their designs. Artists criticise that GenAI may not be able to accurately capture certain nuanced aspects of their planned writing, such as emotional depth and writer identity. Further, Suh et al. (2021) noticed that GenAI’s bounded training scope may limit the extent to which human collabo-rators can contribute diverse expertise outside the GenAI’s domain. The inability of GenAI to understand and illustrate cultural and social contexts that often underpin creative work can lead to outputs that are disconnected from the intended creative vision.

However, Louie et al. (2020) observed that GenAI in particular helped novice users to better express their creative intent and enhance their sense of ownership, self-efficacy in the creative process, and expand their music composition knowledge. Wang et al. (2024) explored which factors influence the intention to use GenAI among artists and found that GenAI lit-eracy and subjective norms positively influenced intention to use GenAI.

4.4.2 Human agency in guiding GenAI creativity

Another theme is the importance of human intervention in working alongside GenAI. While GenAI can generate content, it is the artists’ role to steer the GenAI to ensure the outputs align with their creative vision and the contextual requirements. For instance, Tsao and Nouges (2024) found that students actively intervened to include their own direction and style into the GenAI-generated content. Additionally, artists tend to prefer creative outputs where they have invested more effort, likely due to feelings of ownership, competence, or effort justification. Jansen and Sklar (2021) found through qualitative interviews that artists generally prefer co-creative GenAI and remain sceptical about fully automating creative work.

In line with this, Zhou and Lee (2024) noted that while GenAI can initially boost crea-tive productivity, the novelty of GenAI-generated content decreased as the process became more automated and standardised. Particularly, artists emphasized the need to preserve traditional drawing practices and insist on maintaining editorial control over the final artistic output. Leung et al. (2023) emphasized the importance of co-creation to prevent racism, stereotyping, and various forms of bias and discrimination. However, Bianchi et al. (2023) demonstrated that both user-driven efforts to request diverse representations and institutional attempts to implement protective measures have fallen short in pre-venting content from perpetuating stereotypes. Similarly, Tsao and Nouges (2024) emphasise the importance of human agency in the control of GenAI.

Artists felt that they had to actively shape the content generated by GenAI through specific prompts in order to maintain control over the creative process. GenAI executes the artist’s prompts while requiring human supervi-sion and guidance, shifting the role of the artist from creator to curator of GenAI-generated outcomes. Also, in addition to technical knowledge, the user interface often plays a crucial role in utilising the creative potential of GenAI.

4.4.3 Developing GenAI literacy

Finally, the literature emphasizes the necessity of developing new skills and vocabu-laries for effectively interacting with GenAI. In architecture education, for example, Paananen et al. (2023) suggest that students need to be equipped with a detailed tech-nical vocabulary for describing design concepts for GenAI as well as an understand-ing of the limitations and trade-offs of these technologies compared to traditional methods. Similarly, Oppenlaender (2022) explains the importance of prompt engi-neering and the role of online communities to increase the skills of artists. Kim (2024) demonstrates a significant correlation between language proficiency and the quality of images generated by GenAI tools. Additionally, the growing ecosystem of tools and resources can contribute to supporting creativity in text-to-image gen-eration.

This hybrid approach of blending traditional expertise with GenAI literacy may define the next generation of creative professionals. The propositions derived from the SLRs are depicted in Fig. 4.

4.5 What measurements are used to evaluate content generated by GenAI?

While all of the selected articles are set in the general context of creativity or a creative domain, they employ varying methods for evaluating the content gener-ated by GenAI. Objective and subjective measures are used to evaluate content cre-ated by GenAI within creative contexts. In particular, most of the identified studies used subjective measures to understand individuals’ perceptions of creativity based on broad or specific aspects. Gangadharbatla (2022) used a semantic differential scale to measure attitudes toward artwork and evaluations of originality, creativ-ity, expressiveness, and emotional connection. Notably, the evaluation of artwork differs across various studies, with some measures overlapping and being slightly modified, as seen in Mateja et al. (2024) and Gangadharbatla (2022).

While some researchers focused more on the attitude toward the artwork (see e.g., Demmer et al. 2023; Gangadharbatla 2022; Latikka et al. 2023), others assessed its value in terms of emotional, social, and quality dimensions (see e.g., Mehler et al. 2024; Tigre Moura et al. 2023). Yang et al. (2024) asked how engaging the GenAI-generated content was considered to assess perceived quality on a 7-point Likert scale. In con-trast, Hong et al. (2022) used a 9-item scale to assess the quality of the content, focusing on specific quality attributes of the output in the particular art domain. Tak-ing a different perspective, Badura et al. (2022) utilised emotion analysis and word percentages to quantify the perceived quality of GenAI-generated content.

Mehler et al. (2024) assessed quality, emotional value, value-for-money, and social value, as well as appreciation of AI-generated output using 7-point Likert scales. Zhang and Gosline (2023) assessed satisfaction and interest in AI-generated advertising content using 7-point Likert scales. Other researchers delved into psychological and philo-sophical factors related to aesthetic evaluations. Samo and Highhouse (2023) utilised the Art Reception Survey (ARS) to measure aesthetic judgment and appreciation, covering cognitive stimu-lation, negative emotionality, expertise, self-reference, artistic quality, and positive attraction. Hong et al. (2021) employed the expectancy violation scale to assess two key aspects: firstly, the extent to which the content created by GenAI deviated from participants’ expectations, and secondly, the subsequent evaluation of the content itself.

Several studies focused on specific aspects of creativity and creative expres-sion. Messer (2024) evaluated creative authenticity and novelty with 7-point Lik-ert scales with participants rating the extent to which an AI-generated collection reflected true inspiration and unusualness. Pellas (2023) explored creative identity, asking participants to reflect on their perceived expression of creativity for academic writing. Louie et al. (2020) used a 7-point Likert scale to measure creative expres-sion, self-efficacy, effort, engagement, learning, completeness, and uniqueness in music composition tasks. Magni et al. (2024) employed 5-point semantic differen-tial scales that ranged from “very uncreative” to “very creative” to assess creativity.

In contrast, Paananen et al. (2023) applied the Creativity Support Index (CSI) to explore how GenAI facilitated creative output, with participants rating statements on a scale from “highly disagree” to “highly agree” across dimensions such as collabo-ration, enjoyment, exploration, expressiveness, immersion, and results worth effort. Lastly, Lyu et al. (2023) utilised a qualitative method, using open-ended questions to assess the degree of technical presentation, meaningful connection, and emotional touch in Chinese cultural jewellery designs.

Taken together, a consistent pattern emerged throughout the analysed articles: Researchers viewed creativity as a multi-layered phenomenon and included different dimensions to capture the complexity of creativity. In particular, studies frequently use varying scales to evaluate GenAI-generated content, including novelty, appreciation, and satisfaction. In addition, the evaluation of quality emerged as a recurring theme in several studies. Further, semantic differential scales provide nuanced insights into emotional and expressive dimensions, comparatively Likert scales were widely used for assessing subjective attributes like appreciation or novelty. Additionally, tools like the CSI extend this understanding by capturing immer-sion and collaboration, underscoring GenAI’s facilitation of creativity as a process rather than just an outcome.

Table 3 provides an overview of the measurements used to assess GenAI-generated content.

5 Framework development

The integrative framework in Fig. 5 is derived from a synthesis of the findings of this systematic literature review on GenAI in creative contexts. The thematic analysis revealed the topic’s multidimensional nature. While the framework does not include every part of this complex field, it highlights the most salient factors and interac-tions identified through the systematic literature review. At first, this model shows that artists usually follow a three-stage process to create content with GenAI. This process begins with the artist’s initial input and the AI’s gener-ated alternatives, progresses through a stage of refinement based on artist feedback, and culminates in a phase of mutual learning and alignment.

The interaction between the artist’s skill set and the alignment of creativity plays a pivotal role in each stage of this process, as the artist’s expertise directly influences the quality and direction of GenAI’s outputs. For instance, in the initial input phase, skilled artists are better equipped to craft precise and creative prompts that communicate their vision effectively to the GenAI system. For example, a professional graphic designer famil-iar with prompt engineering can use specific stylistic terms or technical descriptions that result in outputs closely aligned with their intended aesthetic goals. By contrast, less experienced users might provide generic prompts, leading to broader or less relevant outputs that require significant refinement.

However, researchers have put forward that the use of GenAI for novice users can act as a learning tool, helping them to explore creative possibilities, improve their technical skills, and gain confidence in their creative abilities through iterative feedback and experimentation.

The framework identifies several key artwork characteristics that emerge from this systematic review. These include the label of origin (whether the work is labeled as human-made or GenAI-generated), the context of the artwork (commercial or non-commercial), its tangibility, the level of human effort involved, and the type of artwork (abstract or representational). These characteristics influence how the artwork is perceived and evaluated by audiences. Additionally, the framework acknowledges the role of indi-vidual differences in shaping the reception and evaluation of GenAI-generated art. Factors such as gender, personal beliefs about GenAI, art expertise, GenAI literacy, and overall compe-tence are presented as potential moderators. The consequences of this process can be divided into two main categories: evaluation and behavioural intention.

The evaluation aspect encompasses a range of assessments, including authenticity, novelty, creativity, originality, quality, and emotional response. These factors collectively influence the overall perception of GenAI-generated artwork. The behavioural intention category focuses on the future GenAI usage. By mapping out the relationships between content generation processes, artwork characteristics, individual differences, and consequential outcomes, the

Table 3 Summary of measurements for evaluating GenAI-generated content

Table 3 (continued)

Table 3 (continued) framework offers a holistic lens through which to examine the impact of GenAI in creative contexts. For industry practitioners, particularly in fields such as design, and marketing, the framework provides a practical guide for integrating GenAI into their creative workflows. By understanding the factors like “type of artwork” (abstract vs. representational) and “tangibility of artwork” (intangible vs. tangible), organisa-tions can tailor their use of GenAI to align with audience expectations and enhance creative value. For instance, knowing that consumers might perceive less effort in GenAI-generated outputs, companies can strategically disclose the level of human involvement to mitigate potential biases and increase the perceived value of their products.

Further, if a product or artwork is labeled as GenAI-generated, organisa-tions may anticipate scepticism and plan communication strategies that emphasize the collaborative nature of the creation process, showcasing human oversight and creative input to maintain trust and authenticity.

6 Directions for future GenAI research

In this section, the author suggests a number of promising directions for future research that could augment our current understanding of GenAI in creative con-texts and highlights potential gaps for further study for future GenAI research. The author proposes a research agenda to address the identified gaps in the literature, which is presented in Table 4.

6.1 Contextual factors and longitudinal trends

The existing literature suggests a prevalent bias against GenAI-generated art, but this study’s detected research gaps call for further exploration. Current research has primarily focused on conducting controlled experiments in laboratory settings to investigate the evaluation of GenAI-generated content. However, as GenAI becomes more accessible and accepted in the art world, also context will become increasingly important. One can expect that galleries and museums are going to include GenAI-generated art in their exhibitions and collections. To address this gap, researchers could incorporate field experiments to investigate the effects of various contextual factors on the perception and evaluation of GenAI-generated content. Field experi-ments can provide valuable insights that complement the results of current labora-tory studies.

For example, researchers could investigate the perception of GenAI-generated art when it is presented in art galleries or museums. Researchers can observe how these contextual elements influence audience reactions by manipulat-ing framing and curation of GenAI-generated artworks. Similarly, field experiments could examine the performance of GenAI-generated content in commercial contexts such as marketing. This would build on previous research by Messer (2024) and Magni et al. (2024), who found that commercial contexts are perceived differently when compared to non-commercial contexts. Another promising avenue for future research involves the integration of LLMs in the artwork selection process for com-parative studies.

To date, only Köbis and Mossink (2021) have employed LLMs to curate artworks for such comparisons and revealed that this approach significantly influences humans’ ability to detect the artwork’s origin.

With the spread of GenAI and the accompanying integration into creative con-texts, an increase in public awareness of GenAI can be expected. This shift calls for long-term research on how audience perceptions and attitudes towards GenAI-generated content evolve. Lattika et al. (2023) and Zhou and Lee (2024) are the only articles among this study’s selection that have conducted a longitudinal study. Also, while GenAI can enhance productivity and offer new perspectives, there is a risk of over-reliance on these tools. Future studies should examine whether prolonged use leads to creative stagnation or sparks new forms of artistic innovation, addressing Bartlett and Camba’s (2024) concern about over-reliance on these tools and the subsequent risk of less innovative designs.

In addition, Chamberlain et al. (2018) and Gangadharbatla (2022) have pointed out that the genre of art influences an individual’s ability to identify the ori-gin of artwork and affects the evaluation process. However, future research is called for to investigate how this perception changes as technology advances. Differences in artistic styles may become less distinct, and cognitive biases such as associating abstract art with AI and representational art with humans may change over time. Long-term studies could reveal whether advancements in GenAI blur the distinc-tions between abstract and representational art, altering cognitive biases that cur-rently associate abstract art with GenAI and representational art with humans.

Stud-ies indicate that disclosure of GenAI in creative processes often lead to a perception of reduced effort, therefore, it would be interest-ing to investigate whether increased knowledge of working with GenAI can mitigate

Table 4 Future research directions and illustrative research questions this effect and whether this bias changes over time. Another overlooked factor is how the integration of GenAI into creative industries is reshaping the job market by transforming traditional roles and creating new hybrid positions that combine artistic expertise with GenAI literacy. Long-term studies are crucial to understand how GenAI impacts job opportunities, industry standards, and the socioeconomic dynamics of the creative sector, ensuring a balance between automation and the preservation of human creativity. Therefore, researchers should address these topics in future research.

6.2 Human characteristics and individual differences

The current body of research on GenAI lacks an examination of how individual differences and personality traits influence user interaction with and perception of AI-generated content. A promising direction for future studies is investigating whether personality dimensions, such as the Big Five personality traits or the Dark triad, play a role in shaping atti-tudes towards GenAI. The impact of individual differences such as usage frequency, age, gender, and expertise on the perception and application of GenAI needs fur-ther investigation. Understanding how these variables influence the evaluation of GenAI-generated content and how individuals utilise GenAI could offer valuable insights. Further, the existing literature has predominantly focused on students.

Expanding this research to include a more diverse range of users, including professional artists and designers, would be necessary in developing a comprehensive understanding of how

GenAI can be optimally integrated into various creative processes. Another per-spective would be to research the role of cultural background in shaping the evalu-ation of GenAI-generated content and the adoption of GenAI. Recent research on GenAI in creative contexts has predominantly focused on Western perspectives, as evidenced by studies from Gangadharbatla (2022), Koivisto and Grassini (2023), and Millet et al. (2023). However, as AI-driven creative tools gain global traction, there’s a need to broaden our research horizons. While some scholars have begun to explore GenAI’s creative applications in non-Western contexts, these studies remain scarce. This gap is particularly strik-ing given the significant growth in GenAI worldwide, including emerging countries.

Also, research indicates a significant imbalance in cultural representation within GenAI systems, with a pronounced bias towards Western perspectives and imagery. Drawing on prior literature on intelligent robots, westerners have often viewed robots with fear and ethical concerns, influ-enced by narratives such as the ‘Frankenstein syndrome’, which emphasises the dif-ference between humans and machines. In contrast, Japanese people generally see robots as harmonious companions that are characterised by cultural stories without questioning the uniqueness of humans. Further, presenting technology as a means to enhance personal uniqueness can positively influence individual-ists’ attitudes towards AI.

For example, individuals with indi-vidualistic traits are more receptive to personalised suggestions and willing to invest more in AI-driven recommendations that help them stand out from others compared to individuals with collectivistic traits. For collectivist individuals, focusing on altruistic goals can encourage self-disclosure in AI systems, whereas for individual-ist individuals, highlighting self-interest goals may promote self-disclosure. These cultural distinctions not only shape attitudes towards AI but also influence the underlying mechanisms driving technology adoption among different user groups. For instance, effort expectancy emerged as the strongest predictor of ChatGPT adoption in the UK, meaning UK students valued the ease of use of the technology more highly.

In contrast, for Nepalese students, performance expectancy had a more significant impact, suggesting a stronger focus on the perceived benefits of using ChatGPT. Comparative studies across different cultural contexts could reveal how cultural norms, values, and communication styles affect users’ engagement with and acceptance of GenAI-generated content.

Another promising avenue for future research is to explore the impact of GenAI on designers’ sense of ownership and identity throughout the creative process. The sense of ownership is fundamental to designers’ engagement and satisfaction with their work and, therefore, warrants further research. Future stud-ies could examine whether GenAI enhances or diminishes designers’ feelings of attachment and responsibility toward their creations. El-Zanfaly et al. (2023) sug-gest that understanding the interface is more important than understanding techni-cal details. Consequently, researchers can explore how individual differences inter-act with specific features of GenAI tools, potentially leading to more tailored tool interfaces.

Although research by Wang et al. (2024) has identified GenAI literacy and subjective norms as positive influences on artists’ intentions to use GenAI, there remains a significant gap in our understanding of behavioural intentions in creative contexts and the factors that influence them.

6.3 Mitigating bias and stereotyping in GenAI

For future studies, researchers should also investigate how to prevent stereotyping and biases in GenAI contexts. Mitigating stereotyping warrants research attention since there has been a lack of empirical evidence. Previous research has shown that biases in GenAI can lead to negative outcomes, including perpetuating racial ste-reotypes. Researchers that want to address biases and stereotyping in GenAI need to follow a multifaceted approach includ-ing both human agency and technology. For human agency, researchers can look at how to promote diversity and inclusivity in AI development teams. This diver-sity can help identify potential biases that might otherwise go unnoticed. Addition-ally, education and training for AI developers and users on recognising and mitigat-ing biases is necessary.

On the technical side, researchers can apply advanced debiasing techniques such as adversarial debiasing, where a separate model is trained to detect biases in the main model’s outputs, or counterfactual data augmentation, which involves generating synthetic data to balance underrepresented groups in the training set. Developing new tools to make GenAI processes more transparent and interpretable would allow for easier identifi-cation of potential sources of bias. Researchers should explore the ethical implica-tions of these biases and propose guidelines for creating more inclusive GenAI in creative contexts.

6.4 From decision‐making to creativity

A promising avenue for future research involves exploring whether factors that hold for AI acceptance also apply to GenAI. While research on AI has been extensive, there is a gap in understanding whether the identified mechanisms from decision-making domains can be transferred to the context of creativity and GenAI. Burton et al. (2020) and Jussupow et al. (2020) studies on algorithm aversion provide a valuable founda-tion upon which to build. Their findings shed light on how algorithm characteristics and human characteristics shape perceptions and acceptance of algorithmic decision-making tools. Researchers can focus on replicating studies to test whether the same principles hold when we consider the application of GenAI in creative contexts, where subjective preferences, emotional responses, and artistic expression play a pivotal role.

Researchers can test the generalizability of algorithm aversion factors, such as per-ceived capabilities and algorithm performance, to better understand how people interact with GenAI in creative contexts. Bridging the gap between decision-making and crea-tive contexts would advance our theoretical knowledge and inform the development of GenAI technologies. Existing studies on AI acceptance in decision-making use frame-works such as the Unified Theory of Acceptance and Use of Technology (UTAUT), the Technology Acceptance Model (TAM), and the Stimulus-Organism-Response (S-O-R) model, but there is a clear gap in the application of these theories to GenAI in crea-tive domains.

While Cardon et al. (2023) applied the UTAUT model to demonstrate its relevance for GenAI in creative contexts, further research is needed to apply other vali-dated IS theories to address this gap and test their applicability. For exploring the evalu-ation of GenAI-generated content, several established theoretical frameworks may offer valuable insights, such as the framing theory, source credibility the-ory, and attribution theory. These theories could provide robust foundations for understanding how individuals perceive and interpret GenAI-generated artworks. However, there is a lack of application in current research. Additionally, theories such as social contagion theory and social identity theory present promising avenues for exploration in this context. These theoretical frameworks should be applied in future studies on GenAI-generated content.

7 Conclusion and discussion

This literature review has highlighted the impact of GenAI on artists and its observ-ers in creative contexts. This analysis has highlighted several directions for future research that offer potential for theoretical expansion and empirical investigation in this rapidly evolving field. The author anticipates that the issues and questions raised here will be further explored, scrutinised and refined and should ultimately contrib-ute to a more nuanced understanding of this transformative technological paradigm for creative contexts. This study contributes significantly to the understanding of both the consequences and facilitators affecting creative industries in their integra-tion of GenAI. This study presents a conceptual framework grounded in empirical evidence.

To the author’s knowledge, this represents the most holistic research of current literature in the use of GenAI in creative contexts, employing a systematic review methodology.

While machine learning processes are increasingly demonstrating creative capabilities, they may be viewed as complementary tools rather than replace-ments for human creativity. These GenAI systems, dependent on human-gen-erated data for their function, offer new avenues for inspiration and ideation. Rather than replacing artists, GenAI is evolving into a collaborative partner that enhances human creative potential but remains reliant on the original creative output and input of artists. Particularly for emerging artists and novices, GenAI offers unprecedented opportunities to explore and learn, which can accelerate their artistic development and broaden their creative horizons.

Recent studies have shown that individuals’ ability to distinguish between GenAI-generated and human-made content is far from consistent and even experts in creative fields sometimes struggle to differentiate GenAI-generated work from that pro-duced by human artists. This variability stems from a complex interplay of fac-tors, including for instance, the observer’s expertise, the level of human effort, and specific characteristics of the creative output.

The current study is subject to several limitations that should be acknowl-edged. First, the literature search for this study was confined to the citation data-bases Web of Science and ScienceDirect. Although these databases are known for comprehensive documentation of high-quality published literature, the choice of databases may have led to the exclusion of relevant articles that are indexed elsewhere. Similarly, the selection of keywords and search terms, despite efforts to include synonyms and extend the search, may have led to the omission of potentially relevant studies. This keyword and search term limitation is an inherent challenge in systematic reviews, where the compre-hensiveness of the search is constrained by the predefined search strategy.

Finally, although the study employed a rigorous selection process following Tranfield et al. (2003) and Crossan and Apaydin (2010), the process of selecting and cod-ing articles may not be entirely free from subjectivity. Considering that only one author worked on the topic, despite discussing with colleagues, decisions made during the selection and coding stages may have been influenced by personal biases. This subjectivity in article selection and coding is a limitation that should be considered when interpreting the results of this study. While the sample size of 64 studies may appear modest, it reflects the nascent stage of research on GenAI in creative contexts. The author aimed to rigorously filter and select the most rel-evant and high-quality articles to ensure the robustness of this SLR.

Nonetheless, the limited sample size may constrain the generalizability of the findings, under-scoring the need for future studies to expand and build upon this work as the field continues to evolve. Another potential limitation of this review is the inherent language bias resulting from the inclusion of studies published only in English and German. While these languages were chosen due to the author’s proficiency and their prevalence in scholarly literature on the topic, this criterion may have excluded significant contributions from studies published in other languages. Researchers are encouraged to expand upon the insights uncovered in this review, by analysing emergent studies and investigating additional ways in which GenAI may influence the field of creativity going forward.

Appendix.

See Table 5.

Table 5 Overview of the selected articles through the SLR

Table 5 (continued)

Table 5 (continued)

Table 5 (continued)

Funding Open Access funding enabled and organized by Projekt DEAL. The COMET project is funded by the Federal Ministry for Economic Affairs and Climate Action (BMWK) as part of the technology program “SmartLivingNEXT–Artifical intelligence fo sustainable living and residential environments”.

Data availability I do not analyse or generate any datasets, because this work proceeds within a theoreti-cal and mathematical approach. One can obtain the relevant materials from the references below.

Declarations

Conflict of interest The author has no relevant financial or non-financial interests to disclose.

Download transcript ↗