Design principles for text-to-image generative artificial intelligence creativity support tools for visual design
1 More Paper · Full Reading

About this paper
A full audio edition of this paper.
Authors: Savindu Herath, Amirsiavosh Bashardoust, Yonah Bole, Yash Raj Shrestha
Published in: European Journal of Information Systems
Publication date: 2026-01-20
Read the paper: https://doi.org/10.1080/0960085x.2026.2616042
Source license: Creative Commons Attribution 4.0 International — https://creativecommons.org/licenses/by/4.0/
The authors and publisher do not sponsor or endorse this recording.
Transcript
You’re listening to “Design principles for text-to-image generative artificial intelligence creativity support tools for visual design,” by Savindu Herath and colleagues. Published in European Journal of Information Systems on January 20, 2026.
ISSN: 0960-085X (Print) 1476-9344 (Online) Journal homepage: the linked source
Design principles for text-to-image generative artificial intelligence creativity support tools for visual design
Savindu Herath, Amirsiavosh Bashardoust, Yonah Bole & Yash Raj Shrestha
To cite this article: Savindu Herath, Amirsiavosh Bashardoust, Yonah Bole & Yash Raj Shrestha (2026) Design principles for text-to-image generative artificial intelligence creativity support tools for visual design, European Journal of Information Systems, 35:4, 679-704, DOI: 10.1080/0960085X.2026.2616042
RESEARCH ARTICLE
Design principles for text-to-image generative artificial intelligence creativity support tools for visual design a, Amirsiavosh Bashardoustb, Yonah Boleb and Yash Raj Shresthab Savindu Herath aDepartment of Management, Technology, and Economics, ETH Zurich, Zurich, Switzerland; bFaculty of Business and Economics (HEC), University of Lausanne, Lausanne, Switzerland
ABSTRACT.
Generative AI (GenAI) presents significant opportunities, particularly for creative work like visual design. GenAI can effectively address creative challenges like the “blank page” problem by enabling rapid visual conceptualisation, thereby enhancing productivity in visual design tasks. However, organisations face significant challenges in integrating off-the-shelf, untamed foundation models, whose complex user interfaces are misaligned with production-oriented workflows, exacerbating issues such as AI illiteracy, employee resistance, and job displacement.
To facilitate GenAI integration into creative workflows while addressing these challenges, we conduct an action design research study at Ubisoft, where we develop and introduce a GenAI-enabled creativity support tool (CST)—comprising workflow guidance and simplified user interface on Stable Diffusion models—as a bridging tool in concept art creation processes. Our user studies and evaluations indicate that the artefact increases user acceptance, improves productivity, and adds value to both concept artists and their studios. In addition to detailing the artefact’s design and evaluation, we present seven design principles for text-to-image GenAI CSTs for visual design.
These principles contribute to the academic discourse on human-AI collaboration by filling a critical gap in prescriptive design knowledge and emphasising the role of GenAI in augmenting, rather than replacing, human skill and ingenuity.
1. Introduction.
The focus of artificial intelligence (AI) research and its business applications has increasingly shifted towards generative AI (GenAI), driven by the rise of powerful foundation models that can be quickly adapted to specialised tasks such as generating text, music, animation, and designs. Unlike discriminative AI, which focuses on classification and prediction, GenAI models learn from underlying data patterns to generate new, creative content. This shift has expanded the application of AI beyond decision-making processes to include creative activities, particularly in visual design applications. GenAI’s ability to produce original visual content, such as concept art, graphic designs, and product prototypes, positions it as a transformative tool for businesses engaged in visual design and creative industries.
GenAI’s disruptive potential is evident from the explosion of AI-assisted innovation in creative work across domains such as advertising, game development, and filmmaking, driven by image generators like DALL-E 2 and Stable Diffusion (SD), which promise significant improvements in creative
ARTICLE HISTORY
Received 1 July 2024 Accepted 8 January 2026
Although various GenAI tools have demonstrated strong performance in enhancing individual creativity, productivity, work quality, and cost efficiency, organisations encounter substantial challenges when attempting to integrate off-the-shelf, untamed foundation models, with user interfaces (UIs) having legions of configuration options into production-oriented workflows. We identify three prevalent challenges: a) AI illiteracy and skill gaps, b) employee resistance, and c) errors and hallucinations by GenAI. First, GenAI technology progresses more rapidly than the workforce can upskill, leading to AI illiteracy. A skill gap arises when designers lack the requisite knowledge to effectively translate their creative intentions into prompts for successfully operating these tools (Z. Wang et al., 2024), and many organisations struggle to bridge this gap.
AI illiteracy can result in employee resistance or inefficiency, particularly when designers are unfamiliar with how to interact with GenAI sys tems, interpret their outputs, or adjust models for specific requirements, ultimately limiting their ability to fully leverage the capabilities of GenAI. Furthermore, AI tools are frequently perceived as overly complex for everyday use by non-technical staff. Designers often find it difficult to incorporate these tools into their workflows without guidance. Second, the integration of GenAI into visual design processes often triggers resistance from employees, who may fear that these tools will lead to deskilling, undermine their creative agency, and ultimately replace them. Employee resistance may also arise from a lack of trust or confidence in AI-generated outputs, which has been widely recognised as a key contributor to algorithm aversion.
This aversion to AI is particularly pronounced among skilled and expert workers, who are often more reluctant to rely on algorithmic outputs. Third, GenAI systems may generate erroneous content and fail to capture user intention reliably (Z. Wang et al., 2024), contributing to the perception that these models lack stability. This raises concerns that the outputs could lead to undesirable or incorrect practices and, more critically, could diminish human autonomy and creativity in the design process.
Existing theoretical frameworks on AI design provide limited prescriptive guidance for the design of GenAI artefacts capable of overcoming these challenges in practice, given that GenAI represents a novel and fundamentally different phenomenon from discriminative AI. Despite the increasing demand for actionable design knowledge to effectively integrate GenAI into creative processes, the current research in this domain remains predominantly limited to controlled laboratory experiments or highly technical approaches and lacks practical, application-oriented solutions.
Extant design science research offers inadequate guidance regarding GenAI’s integration into existing organisational processes with context-specific demands, its component (model and UI) and interaction (prompting) design, and how best to design or customise these components in a way that ensures user acceptance and enhances productivity rather than replacing humans or deskilling the workforce. The dearth of prescriptive knowledge on designing GenAI artefacts within organisations represents a significant lacuna in the information systems (IS) field. Design principles (DPs) are essential for disseminating prescriptive design knowledge to broader audiences. Effectively formulated DPs should include three types of information, action potentials, material properties, and boundary conditions. Action potentials refer to the actions enabled by the artefact.
Material properties pertain to the characteristics of the artefact that support these actions. Boundary conditions specify the contexts in which the design operates effectively. Notably, the absence of prescriptive design knowledge on the material properties of GenAI artefacts, the action potentials these properties enable the users, and their boundary conditions poses challenges for practitioners in navigating the complexities of integrating GenAI into organisational processes. This limits the potential for fostering effective collaboration between humans and GenAI in organisations.
Recognising these practical and theoretical challenges, this study explores how to design and integrate text-to-image GenAI artefacts into visual design processes with the goal of developing design knowledge in the form of actionable DPs. The visual design context was chosen due to its significance as a key application area of GenAI. To address the research question, we conducted an action design research (ADR; Sein et al., 2011) study at Ubisoft (the linked source), a leading European game development company.1 Game development relies heavily on a wide range of visual design processes, and according to Bain and Company’s report, GenAI’s contribution to video game content may grow from less than 5% today to 50% in the next five to 10 years.
Within Ubisoft, we develop and introduce a text-to-image GenAI creativity support tool (CST)—comprising a guidance workflow and simplified UI on SD2 models—in concept art creation processes. More specifically, our artefact aimed to effectively address creative challenges like the “blank page” problem which can significantly hinder workflow efficiency and innovation by enabling rapid visual conceptualisation and translating creative user intentions to optimised prompts, thereby reliably producing visual outputs and enhancing productivity in concept art creation. We draw on a rich set of primary data for the design and development of the artefact.
This includes qualitative and quantitative data; field observations; multiple experiments with GenAI models to compare prompts, model parameters, and generated outputs; internal task-based user studies with six participants (P1 to P6) from Ubisoft, including artists and technical directors, followed by a Qualtrics survey and semi-structured interviews; and an external user study with six participants followed by semi-structured interviews. This diverse set of data and participants provided insights into the effectiveness and usability of the developed artefact, offering a comprehensive understanding of the artefact’s impact on the visual design process.
Our main contribution is seven DPs that provide prescriptive knowledge on developing process guidance workflows, prompt engineering, model parameters, output performance metrics, and UI. These DPs encapsulate the material properties of a GenAI artefact for visual design processes, the action potential these artefact properties offer to its users, and the boundary conditions. The DPs emphasise increasing user acceptance and practitioner relevance, prioritising augmenting rather than automating the visual design processes, fostering co-learning opportunities between users and GenAI, minimising errors and hallucinations, and designing against deskilling and the replacement of the workforce, which are all viewed as pressing challenges in contemporary human-AI collaboration scholarship and practice.
Our principles contribute to IS scholarship on GenAI-assisted visual design by providing guidance on harnessing the potential of GenAI to overcome creative blocks and enhance productivity, thereby contributing to the field of human-AI collaboration. The principles are expected to benefit and guide practitioners in developing customised text-to-image GenAI CSTs in other creative fields with similar settings, such as filmmaking, fashion design, and animation.
2. Theoretical background.
2.1. Human-AI collaboration.
In the evolving landscape of organisational operations, the integration of AI has transformed the roles of humans in decision-making and problem-solving processes. This transformation of organisational operations has primarily manifested in two forms: automation and augmentation. Automation refers to AI fully taking over a human task, whereas augmentation involves collaboration between humans and AI in task execution. Scholars widely agree that AI is poised to replace humans in performing tasks that are routine and easily automated. Typically characterised by their repetitive and structured nature, these tasks are ideal for AI, which can be trained on extensive historical data to surpass human performance.
Through effective automation of routine tasks, organisations can achieve improvements in speed and efficiency, consistency, and accuracy, while freeing up human resources for more critical, complex tasks. However, the widespread adoption of AI has reignited concerns about job displacement resulting from increased automation.
While automation offers clear advantages in handling routine tasks, there is some debate about the extent to which AI can assume roles involving complex cognitive tasks. For tasks that demand adaptability, flexibility, creativity, and personal judgement, human workers are still considered irreplaceable. For tasks such as managerial decision making, scholars have advocated for augmentation through effective human-AI collaboration. Augmentation is based on the complementary strengths of humans and AI. AI excels in information processing and predictive capabilities, offering consistency, speed, and efficiency, as well as reducing cognitive biases, whereas humans contribute domain expertise, contextual understanding, and the ability to navigate unforeseen circumstances.
These complementary strengths enable effective human-AI collaboration, leading to superior performance in complex cognitive tasks. Numerous frameworks for human-AI collaboration, such as human-in-the-loop systems and human-AI ensembles, have been proposed to effectively integrate the complementary strengths of humans and AI.
The discourse around human-AI collaboration is enriched by the development of GenAI technologies and their applications across various domains, including text and image generation, as we move beyond predictions and decision making towards creative problem solving. Foundation models like GPT-3, LaMDA, and Wu-Dao have paved the way for innovations such as DALL-E, ChatGPT, and various generative applications, indicating a shift towards more collaborative and creative interactions between humans and AI. Studies have examined the dynamics of hybrid problem-solving processes and shown that GenAI can enhance human capacity for creative problem solving in settings like call centres. Human-AI collaboration has been shown to have the potential to increase creativity in the workplace, particularly among higher-skilled employees.
In visual design tasks like art creation, Zhou and Lee (2024) have shown that artists skilled at generating innovative ideas before using GenAI are judged more positively if they continue to effectively explore novel concepts after adoption, particularly if they can adeptly handle and organise them into coherent artworks. In fashion design, Davis et al. (2024) found that GenAI tools facilitate divergent and convergent thinking, where divergent thinking involves generating multiple, varied ideas, and convergent thinking refers to narrowing down and refining those ideas in design space exploration. Hence, the rise of GenAI offers organisations the potential to harness greater benefits, including enhanced creativity, efficiency, and productivity.
2.2. Impact of GenAI on productivity.
One of the most promising advantages of human-GenAI collaboration is the enhancement of employee productivity in creative tasks. Both controlled experiments and field studies offer compelling evidence across various domains that GenAI significantly improves productivity in guided tasks. In a field study involving software developers, Ziegler et al. (2024) observed notable improvements in task performance and accomplishment with the integration of GitHub Copilot into development workflows. Similarly, in an experiment with customer service agents, Brynjolfsson et al. (2025) reported a 14% average increase in productivity, measured as issues resolved per hour, with the use of a GenAI-powered chat assistant. This particularly benefited low-skilled workers, who saw an impressive 34% productivity boost.
In an online experiment with college-educated writers, Noy and Zhang (2023) found that access to ChatGPT reduced task completion time by 40%, while also improving output quality by 18%.
Despite the strong evidence of GenAI’s positive impact on individual productivity, its effect at different skill levels is more nuanced. Interestingly, low-skilled or less experienced employees tend to experience disproportionately higher productivity gains compared to their more skilled and experienced counterparts when using GenAI. The automation and augmentation capabilities provided by GenAI tools enable workers with varying levels of skill and experience to perform similar tasks more effectively. For instance, in coding, GenAI assists less-experienced programmers in producing efficient, clean code, while in fashion design, it supports novice designers in exploring design spaces and generating innovative design concepts, enabling them to achieve results comparable to more experienced professionals.
However, emerging research indicates that GenAI tools may also widen the productivity and skill gap among creative professionals. These studies indicate that the creative enhancement provided by GenAI is not uniformly distributed, but rather dependent on individual characteristics such as skill level, experience, and creative style (e.g., exploratory vs. non-exploratory). For instance, Jia et al. (2024) found that GenAI assistants boost the creativity and productivity of higher-skilled employees to a much greater extent than that of their lower-skilled counterparts, thereby widening the productivity and skill gap. This growing divide leads to negative emotional experiences for lower-skilled employees, diminishing the overall benefits of GenAI for both creative professionals and their organisations.
Consequently, the design of GenAI tools becomes critically important: These tools should reduce rather than exacerbate the productivity and skill gap. In creative fields such as visual design, CSTs are particularly crucial for helping workers overcome creative blocks and enhance their productivity.
2.3. Creativity support tools.
CSTs are defined as digital systems that incorporate creativity-focused features to positively influence users with varying levels of expertise across one or more phases of the creative process. Research on CSTs has addressed both abstract questions, such as how general-purpose tools and technologies can foster creativity, and the development of specialised tools designed to support distinct aspects of creative workflows. By doing so, the CST literature has made valuable contributions that aid creative professionals in overcoming creative blocks at different stages of the creative process. For instance, some CSTs assist in the ideation phase by facilitating rapid idea generation, effectively solving the “blank page” problem, while others support both divergent and convergent thinking during design space exploration.
CSTs vary widely in terms of their target audience, their point of impact within the creative process, their level of intrusiveness, and the types of technologies employed. As technology has advanced, the field of CSTs has evolved accordingly. With the rise of GenAI in recent years, researchers have increasingly focused on the potential of GenAI in CSTs, applying it to areas such as 3D scene design, fashion, storytelling, art creation, and character development. These tools are designed to enhance the creative process by leveraging GenAI models, while ensuring that human agency is preserved and fostering collaborative interactions.
In the realm of visual design, for example, Oh et al. (2024) introduced LumiMood, a GenAI-based CST that automatically adjusts lighting and post-processing to create desired moods for 3D scenes. Similarly, Davis et al. (2024) developed a GenAI tool for fashion design that supports both divergent and convergent thinking. For narrative creation, Chung et al. (2022) presented TaleBrush, which uses line sketching interactions combined with a GPT-based language model to assist in story generation. Extending this concept to character development, Qin et al. (2024) explored CharacterMeet, a system that facilitates iterative character construction through dialogue with chatbot avatars powered by large language models.
As GenAI-enabled CSTs have gained prominence, scholars have examined their impact on productivity. In the context of digital art creation, which closely aligns with our research, the adoption of GenAI-enabled CSTs has been shown to enhance artist productivity by 25% and increase the likelihood of favourable evaluations by 50%. Furthermore, in the ideation phase of visual design, these tools can effectively help creative professionals overcome blocks like the blank page problem by enabling the rapid visual conceptualisation of artistic ideas, thereby generating productivity gains for both individuals and organisations.
However, one of the most pressing functional challenges associated with GenAI-enabled CSTs is the issue of hallucinations—the generation of nonsensical or erroneous content that, while incorrect, appears plausible and is presented with confidence —often compounded by the system’s failure to consistently capture and align with the user’s creative intentions (Z. Wang et al., 2024).
While the productivity benefits of GenAI-enabled CSTs are evident, a deeper examination of how these tools are developed and adapted for organisational contexts reveals a crucial gap in design science research. Existing design science research on AI has predominantly focused on the development and evaluation of decision support systems. This focus is rooted in the capabilities of earlier predictive AI systems, which were primarily limited to assisting decision making by providing analytical insights derived from historical data. However, the emergence of GenAI has broadened the scope of AI applications, extending beyond decision support to include active roles in creative processes, such as visual design.
In practical organisational contexts, companies often require customised GenAI artefacts tailored to their specific industries and operational needs in order to fully leverage the potential of GenAI. Despite this, most research on GenAI remains experimental and centres on generic artefacts, offering limited prescriptive guidance for developing or customising these artefacts for day-to-day operations. This gap underscores the need for prescriptive design knowledge specifically tailored to GenAI-enabled CSTs, which can better guide organisations in effectively integrating GenAI into their visual design workflows.
3. Study context and problem formulation.
3.1. Study context.
The primary objective of this study is to develop prescriptive design knowledge tailored to text-to-image GenAI CSTs in visual design processes. Our focus is on enhancing the efficiency and productivity of creative professionals by addressing creative blocks. We carry out an ADR study at Ubisoft, a leading video game publisher and developer headquartered in Montreuil, France. Ubisoft is known for creating innovative and popular game franchises such as Assassin’s Creed, Far Cry, and Tom Clancy’s. Founded in 1986, Ubisoft has grown into a global entity with around 20,000 employees, operating multiple development studios worldwide.
In our ADR, we focused on concept art creation. Concept art creation involves the production of visual materials that guide the development of video games, serving as a blueprint for the game’s visual and thematic design elements. We focused on human-AI collaboration in the creation of game characters. Concept artists received assistance from GenAI, which generated images in response to prompts fed in by the artists. We explore the interaction between concept artists and GenAI, specifically investigating how GenAI can boost the efficiency and creative productivity of concept artists. Our ADR team consisted of four researchers and a technical director (practitioner), with concept artists serving as the end users.
One researcher was dedicated fulltime to the technical development of the artefact at Ubisoft, ensuring the tool’s alignment with the organisation’s workflows and technical requirements. The remaining researchers played pivotal roles in guiding the technical development, overseeing the artefact evaluation, and shaping the research design and positioning of the study within the broader academic discourse. The technical director provided practical insights from an industry perspective and facilitated collaboration within the company. The concept artists contributed by evaluating the artefact through hands-on usage and participating in focus group discussions and internal user studies, offering critical feedback to refine and enhance the artefact’s usability and functionality.
Following Sein et al. (2011), we conducted ADR in four stages: problem formulation; building, intervention, and evaluation (BIE), which included two sequential design cycles; reflection and learning; and formalisation of learning (see Figure 1).
Table 1 outlines the research activities and sources of empirical data, along with the corresponding research activities and analytical techniques used across the ADR stages, illustrating how empirical insights informed artefact refinement and the derivation of design knowledge.
3.2. Problem formulation.
In game development, concept artists are responsible for designing and creating the initial look, style, and visual feel of a game before their artwork is rendered into two or three dimensions. They sketch and reference key elements like characters, props, weapons, architecture, and environments. The work of concept artists is both significant and demanding due to four key factors. First, developing concept art requires meticulous care and long periods of focused attention from the artists. Their pre-production work establishes a unique artistic direction for each game, enhancing its immersiveness and authenticity. Concept art not only influences the look and feel of a game but also inspires and guides other key contributors, including 3D modellers, animators, and other artists throughout the development process.
Second, concept artists often face significant challenges, including working under tight deadlines, adapting to evolving project requirements, maintaining consistency in style while being adaptable, and effectively communicating their ideas to clients and team members. They must work under intense time pressure to meet launch deadlines, achieve faster time to market, control production costs, and coordinate with various stakeholders to ensure the game’s success. The following statement from one artist emphasises the time-value compromise (P1 Q2; Appendix A):
Creating colour and texture variations for a game object is time consuming and adds little value.
Third, as the gaming industry is highly competitive and game design is heavily influenced by state-of-the-art technologies, artists must navigate creative blocks, handle feedback and criticism, and stay abreast of the latest software and technologies. This is evident in statements from concept artists:
[These days, concept art] is much more technical in computing, whereas before concept art was technical in the way you drew. Before I had to learn to master Photoshop and now I have one more tool [GenAI] to master, which takes time. (P2 Q6)
Finally, concept artists’ work also demands significant coordination. They are expected to work closely with game designers and developers to ensure that the visual elements align with the gameplay mechanics and narrative. The art produced during this phase is crucial for discussions, feedback gathering, and communicating ideas within the development team and to stakeholders, as well as for marketing purposes before a game’s release.
These contextual challenges reflect broader issues commonly associated with the integration of GenAI in creative workflows. First, the steep learning curve associated with mastering tools like GenAI adds to the existing pressure on artists, highlighting a systemic skill gap and AI illiteracy. Second, the fast-paced nature of game development and evolving project demands can intensify concerns around employee resistance, particularly when AI tools are perceived as undermining creative autonomy or replacing manual expertise. Third, artists’ reliance on GenAI for ideation and visual output increases the risk of errors and hallucinations.
Our research was inspired by the important role of concept artists, the practical problems they face, and the potential of GenAI to provide viable solutions. Effectively overcoming creative blocks during the ideation phase and efficiently managing tight deadlines is essential not only for the performance of concept artists, but also for the overall functioning of the game studio. Drawing inspiration from the human-AI collaboration literature and recent research into CSTs, which has proposed GenAI as an effective method for overcoming creative blocks in visual design processes, we aimed to address the challenges concept artists face by developing a novel GenAI artefact (Generative.ConceptCrafter) to facilitate the visual conceptualisation process by efficiently and effectively generating concept art from instructions written in natural language (prompts).
By facilitating fast ideation and sketching, reflection, presentation, and feedback gathering, the GenAI CST Generative. ConceptCrafter streamlines the concept art creation process in game development, ultimately enhancing efficiency and productivity.
Generative.ConceptCrafter is a CST powered by SD models—open-source, text-to-image GenAI models developed by Stability AI.3 The development of the artefact began with off-the-shelf SD models and the WebUI interface. Over the course of our ADR study, we applied a user-centred design approach to expand and customise the tool, transforming it into a full-fledged CST tailored to meet the specific organisational needs and address the practical challenges identified in the introduction. Generative.ConceptCrafter consists of three key components. The first component is a streamlined concept art generation process guidance workflow for the artists to follow when using the standard SD WebUI. This process guidance workflow offers guidance to concept artists in navigating the intricacies of the standard SD WebUI, prompt engineering, and model parameters.
The second component is a WebUI (later customised to ArtistUI), serving as the interface between the artist’s intent and the AI-generated image. The third component is the SD image generation model, which operates in the backend (see Figure 2). Generative.ConceptCrafter, which is an instantiation of text-to-image GenAI CST, aimed to assist concept artists to overcome creative blocks through rapid visual conceptualisation of diverse artistic ideas, ultimately improving their productivity in concept art creation.
4. Artefact development.
The BIE stage consists of two design cycles (see Figure 1). During Cycle 1, we concentrated on prompt engineering and the configuration of GenAI model parameters, leading to the development of the concept art generation workflow and its subsequent user evaluation. Based on the feedback obtained in Cycle 1, Cycle 2 emphasised the development of multiple iterations of a user-friendly interface (ArtistUI), followed by comprehensive user evaluations to refine the whole system.
4.1. Prompt engineering and GenAI model.
parameters
In this section, we outline the extensive testing conducted to identify the optimal prompting techniques and GenAI model parameters, with the goal of establishing the process guidance workflow for image generation using the standard SD UI. Prior research has underscored that prompting techniques and the choice of GenAI model parameters are crucial in conveying the artist’s intentions to guide the GenAI models in the backend (interaction design; Giray, 2023; Liu & Chilton, 2022). Accordingly, we began artefact development by systematically exploring these two design elements. The findings informed the creation of the process guidance workflow (see Section 4.2).
4.1.1. Prompt engineering.
Prompt engineering is the process of structuring instructions in such a way that they can be understood and interpreted by GenAI models to yield desired results. Prompting initiates the image generation process, with users providing written instructions to guide GenAI models in producing desired outputs. To receive expected and controlled results, users need to describe various aspects of the intended output image, such as the subject, its context, lighting, colours, artistic style, perspective, ambiance, etc. Prior studies have highlighted that users often struggle to articulate their creative intentions as effective prompts, a challenge we also identify as the intention-to-prompt gap (Z. Wang et al., 2024). Prompt engineering is a relatively new skill that professionals working with GenAI solutions require, and it is swiftly spreading across various industries and tasks.
Prompting is key for interaction design. Several important aspects constitute a high-quality prompt. First, it should comprise relevant and specific keywords that can be organised into distinct categories (see Table 2). In our tests (see Appendix B), we found that the order of these categories is relevant for effective prompting. Second, GenAI models apply rules to determine what to emphasise in the prompt. For SD, the first keywords of the prompt have more weight than the last ones. Users can manually tweak a keyword’s weight using brackets followed by a multiplication factor (e.g., (keyword:1.1)) increases the weight of keyword by a factor of 1.1. With changes in weights, one can observe variations in the GenAI’s output and apply a sequence of changes until the appropriate configuration is found.
Third, we created presets (Appendix C) containing keywords that can be used as templates for short prompts and negative prompts to add to the user’s written prompt. These templates help users to configure the preset prompt in the artefact to create comprehensive prompts. While prompts are used to generate the desired output, negative prompts are used to avoid or exclude certain elements, ideas, or characteristics leading to the generation of more controlled images by contrasting with the regular prompt. Negative prompts support error reduction and minimise hallucinations. Incorporating these elements in a prompt collectively shapes the final generated image, facilitating more controlled and tailored results closer to the artist’s expectations. An example of this process is illustrated in Figure 3. The negative prompt was:
Negative prompt: low-quality, watermark, deformed, distorted anatomy, poor artistry, NSFW, cleavage, nudity, naked, explicit content, mutated features, extra limbs, extra fingers, ugly, missing body parts, disconnected elements, malformed, abnormal proportions, aberrant hands, aberrant feet, aberrant legs, aberrant fingers
Our approach to prompt engineering for image generation workflow prioritises consistency, reusability, and adaptability across diverse game development contexts and users. To achieve this, our method consisted of several key steps: 1) categorising prompts into distinct keyword categories, as outlined in Table 2; 2) initiating prompt building as an sequential iterative process, initially focusing on selecting keywords for style, media, and scene; 3) gradually incorporating additional keywords to refine the generated image, assessing the direct impact on the output with each addition; and 4) iteratively experimenting and testing different keyword combinations until the optimal prompt configuration was achieved. This systematic approach ensures the development of prompts that are effective and versatile for various applications and user scenarios. An example is provided in Figure 3.
It is important to categorise the keywords, as explained before. Arranging these categories in a sequential order is even more crucial. This became evident when we chose a prompt combining two subjects: a koi fish and a gun—objects that are not often used together. Image generation models tend to dissociate the fish object from the gun object. The results of this experiment are illustrated in Figure 4. The colour coding represents different keyword categories from Table 2.
After extensive experimentation and testing (see Appendix B), we determined that the subject, media, camera, topic, and extra tag sequence yielded the most effective results in our prompt engineering process. Therefore, we decided to adopt this specific order in our workflow to ensure consistency and optimise the performance of our prompting approach.
Our prompt engineering process is designed to bridge the intention-to-prompt gap by refining and optimising the prompts provided by concept artists. This is achieved by categorising keywords, sequencing them in an empirically validated optimal order, and integrating them with negative prompts using presets and templates tailored for game character development. The system supplements user input with predefined prompt categories, default prompts, and additional descriptive tags. This approach enhances the quality and relevance of the generated image while reducing the burden on users to craft lengthy, technically complex prompts.
Moreover, concept artists could emphasise specific keywords in alignment with their creative intentions by adjusting model attention weights, thereby gaining finer control over the generation process. We further extend this framework through multimodal prompt engineering, enabling the use of reference images in conjunction with ControlNet. This allows users to condition the output on spatial structures or stylistic elements present in the reference material (e.g., a sketch), providing a more precise and expressive mechanism for translating creative intentions into prompts and then outputs.
In this section, we focused on identifying the optimal prompting techniques to incorporate into the process guidance workflow, aimed at guiding user interactions with our artefact. By doing so, we enhance user engagement by emphasising the crucial role of human expertise in crafting effective prompts, thereby fostering a more collaborative relationship between artists and the artefact. Next, we experimented with the GenAI model parameters to better understand SD model behaviour—another critical element for developing the process guidance workflow.
4.1.2. GenAI model parameters.
Three key parameters of SD models are crucial for image generation, namely a) the model checkpoint, b) the sampler, and c) the steps. Model checkpoints are pre-trained SD model weights based on a specific SD model version (e.g., SDXL4 v1.0, SD v1.5) for generating images with a particular artistic direction (e.g., specialised in realistic, anime, etc.). In our artefact, we generated images with a realistic, SD v1.5-based model checkpoint for faster image generation (vs. SDXL v1.0) and for its higher compatibility with the extensions required during the workflow, such as ControlNet.5
GenAI images are generated through a denoising process. Samplers define how the image is denoised, and the iterations used for this process are the steps. Generated images require a minimum number of steps to achieve clarity and eliminate visible noise. The noise reduction process initially removes a substantial amount of noise, followed by a gradual refinement to capture finer image details.
The choice of the sampler and steps determines the speed at which the final image is generated. Our artefact employs Euler A or UniPC samplers when we prioritise speed over details and the DPM++ SDE Karras sampler for high-quality final image generation for clearer and more detailed images.
Features that enable artists to craft effective prompts and adjust GenAI model parameters based on their artistic direction and priorities reinforce the perception of GenAI as a collaborator rather than a mere replacement. A comprehensive understanding of how prompt engineering and GenAI model parameters collectively influence the final generated image is essential for establishing the foundation of the concept art generation process guidance workflow.
4.2. Concept art generation process guidance.
workflow
The concept artists in our study lacked technical proficiency in GenAI technologies and thus required guidance in navigating the intricacies of the WebUI, prompt engineering, and model parameters. To address this, we developed the process guidance workflow in the form of a flowchart to guide them, in addition to providing the standard SD WebUI. This workflow was developed following a comprehensive process of experimentation and testing with prompt engineering and model parameter configurations, as we explain and show in the previous sections. Many prompting techniques were tested, we iterated a lot between different ways of prompting, and we ultimately decided to follow prompt categories and a specific sequential order. Model parameters were also tested iteratively.
The results of this extensive experimentation showed that an efficient way to generate images was to first use the parameters for a fast generation process, then enhance and upscale the image in a slower process with a higher resolution output (see Section 4.1 for tests and results).
The resulting workflow encompasses four key steps (Figure 5):
Initial image generation: Using a reference image, users formulate prompts based on existing templates to generate four character views, selecting one from the resultant images.
Upscaling and texture variation: The chosen image undergoes upscaling and texture variation, producing fewer images with high-resolution parameters.
Photoshop sketch and Inpaint: Users upload the preferred image to Photoshop for sketching and realistic character part generation, utilising the inpainting feature.
Final upscale: The image is further upscaled using a specialised extension to enhance facial details, increasing the resolution from 1024 × 476 (step 1) to 2054 × 940 (step 4).
This step-by-step workflow offers valuable insights into the functionalities of specific prompts and parameters, helping artists comprehend the logic behind the generation process. This understanding is particularly crucial for individuals who are unfamiliar with GenAI terminology or the underlying mechanisms, ultimately fostering greater confidence, acceptance, and competence in their interactions with the artefact. After developing this workflow, we implemented it in operational contexts to assist concept artists. This integration enabled artists to interact with the artefact in the organisational environment, facilitating evaluation of its utility and refinement based on received feedback.
4.3. Refinement of the standard SD WebUI into.
ArtistUI
Following the evaluation of the concept art generation workflow with concept artists, it became evident that the original SD WebUI was overly complex, even with the assistance of the workflow (see Section 5.1). Thus, our focus shifted towards sequentially improving the original WebUI over four versions (ArtistUI v1.0 to v1.3; see Appendix D) to develop a more user-friendly interface in Design Cycle 2. As the interface became increasingly customised through iterative development, we renamed it ArtistUI to reflect its adaptation for professional concept artists’ needs. Each updated version of ArtistUI incorporated additional features or automation enhancements in response to user feedback. The first version of ArtistUI (v1.0, Figure 6(a)) automates much of the workflow, simplifying the concept art generation process significantly.
In this version, users are only required to input their prompt, with most parameters being pre-configured for each step. Users enter a short prompt, specify the desired style, camera perspective, artistic media, and negative prompt in addition to the pre-configured prompts visible in the debug info “generated image details”. The goal was to reduce the features and parameters from the SD WebUI that the concept artist had to configure. However, this was deemed insufficiently simplified, as specifying the style, perspective, etc. would not be necessary if the UI were designed for a specific use case in which the user could specify the artistic direction, which would then be automatically provided by the UI. ArtistUI v1.1 (Figure 6(b)) simplified the interface further by hiding unnecessary debug information.
ArtistUI v1.2 (Figure 6(c)) introduced the ControlNet extension, providing users with more control over texture variation. This version also included functionality for both image-to-image and text-to-image generation using ControlNet’s most common models (canny, depth, scribble, and line art). An added feature was a pop-up window to choose from text-to-image, image-to-image, or text-to-image in ControlNet and to choose from the configured ControlNet models when running the code. In the final iteration, ArtistUI v1.3 (Figure 6(d)), all four steps of the workflow were integrated directly into the interface. As the ArtistUI evolved, prompting was also streamlined, requiring only the subject input from the user, with the system combining this input with predefined categories specific to the use case, default prompts, and extra tags to enhance the generated image.
This means concept artists no longer have to write long, detailed prompts.
5. Artefact evaluation.
We evaluated our artefact using the Framework for Evaluation in Design Science Research (FEDS; Venable et al., 2016). Our evaluation consisted of four distinct episodes, incorporating both formative and summative assessments across artificial and naturalistic paradigms (see Figure 7 and Table 3).
In the first evaluation episode, we conducted extensive experimentation and simulations involving various prompting and model parameter configurations outside the company’s concept art creation workflows, effectively in a lab setting. The primary objective of this episode was to establish the PoC, demonstrating that the artefact is functional and capable of producing appropriate results that avoid hallucinations and could be beneficial in concept art generation workflows.
For the second, third, and fourth evaluation episodes, our aim was to assess the artefact’s utility, user acceptance, and productivity enhancements for concept artists, thereby establishing the PoV and PoU of the artefact in practice. Our evaluation strategy focused on rigorously establishing that the artefact’s utility, acceptance, and benefits would persist in real-world situations and over the long term, as recommended by Venable et al. (2016).
5.1. Evaluating the concept art generation.
process guidance workflow
Six (P1 to P6) studio members (concept artists, developers, and managers) with diverse backgrounds and varying levels of familiarity with GenAI were chosen to assess the concept art generation workflow. They were asked to complete their routine game character development tasks (see Appendix A) using the workflow as a guide. Following this, they were invited to participate in
Note: Human-AI collaboration and human-computer interaction (HAI/HCI).
an evaluation survey (see Appendix A). The survey requested details about their background, such as their role, experience, familiarity with AI, and previous usage of SD. Additionally, they were asked to provide feedback on their experience using the workflow, particularly focusing on its usability and its effectiveness in enhancing their job tasks.
Although all participants achieved quality results using the provided workflow (Design Cycle 1), the survey unveiled a predominantly negative sentiment towards the SD WebUI workflow (see Table 4). Users expressed confusion about the parameters and found the workflow challenging to follow. One participant recommended simplifying the solution by reducing the number of parameters and offering a single prompt option. The focus group participants complained about the lack of clarity and complexity of the workflow:
Sometimes [the workflow] is not clear for some parameters, it should be better specified. (P3 SQ5)
Some offered suggestions for improvement:
Some screenshots would have helped to better follow each step, or a guided video. (P3 SQ6)
[A solution with] less parameters, single prompt solution. (P4 SQ6)
5.2. Evaluating Generative.Conceptcrafter.
5.2.1. Internal evaluation.
The internal evaluation involved semi-structured interviews with the same studio members at Ubisoft (see Appendix A), focusing on integration, efficiency, creativity, UI experience, comparative analysis, usability, accessibility, and future perspectives. Participant-generated images from the workflow (Section 5.1) and Generative.ConceptCrafter (with ArtistUI, Section 5.2.1) evaluation phases were collected. Figure 8 shows three of the four workflow steps (S) alongside the images generated by participants.
We classify the problems gathered through user feedback about Generative.ConceptCrafter into three dimensions: 1) complexity vs. simplicity, 2) control vs. ease of use, and 3) specificity (narrow search) vs. generality (broader search). Next, we explain the design decisions pertaining to ArtistUI along these dimensions.
5.2.1.1. WebUI to ArtistUI 1.0 to ArtistUI 1.1.
The complexity of the WebUI presents a significant challenge for users, particularly concept artists who may not have extensive technical knowledge. Having to navigate through numerous features and configure various parameters can lead to cognitive overload, hindering users’ ability to effectively utilise the system. This complexity introduces a trade-off between the
Note: Response codes were Strongly disagree, Somewhat disagree, Neither agree nor disagree, Somewhat agree, Strongly agree.
richness of the features and the simplicity of use. While a highly customisable system offers flexibility, it also poses challenges in terms of feature and parameter selection. On the other hand, simplifying the system may enhance usability but risks sacrificing its ability to address diverse user needs and scenarios. Striking the right balance between complexity and simplicity is essential to ensure that the GenAI CST remains both powerful and accessible.
5.2.1.2. ArtistUI 1.1 to 1.2.
Integrating ControlNet functionalities into ArtistUI 1.2 provides users with enhanced control over the generated output. The addition of visualisations aids users in tuning the GenAI more effectively to achieve desired results. By incorporating these functionalities, the GenAI system moves away from being a black box, allowing users to have more input and control over the creative process. This increased transparency and user involvement contributes to the explainability of the art generation process, addressing concerns related to the opacity often associated with AI-generated outputs. However, ensuring that these additional functionalities are seamlessly integrated and intuitively accessible to users still requires a compromise.
5.2.1.3. ArtistUI evolution over four versions.
The use of predefined keyword categories, default prompts, and extra tags in ArtistUI simplifies the interface for the users yet imposes constraints and promotes specificity in the generated outputs. By structuring the input parameters, the system aims to mitigate the risk of hallucination. Streamlining the input process enhances specificity, aligning user inputs with the desired objectives and reducing the likelihood of erroneous outputs. Ensuring that the predefined categories and prompts accurately capture the diversity of user intentions and preferences is essential to the effectiveness of the system.
Compared to the workflow evaluation conducted via survey (Section 5.1), the interviews revealed a more positive response towards the ArtistUI. The primary factor contributing to this positive change in perception lies in the incorporation of user feedback across four iterations of ArtistUI, resulting in a simplified interface that conceals technical intricacies from users, as demonstrated above. We successfully achieved equilibrium across the identified problem dimensions to align with the preferences and requirements of our user base. Participants unanimously praised the increased efficiency and user friendliness for generating character sheets irrespective of their skill and experience levels: The simple interface with a single prompt greatly simplifies the process (P1 Q1).
Users highlighted the benefits of a simplified UI: It’s much easier to use, with far fewer buttons, and it’s very practical (P2 Q4), describing the process as faster, easier (its ease of use (P6 Q7)), and more straightforward compared to the standard SD WebUI guided by the workflow. Users also affirmed the tool’s usefulness in reducing the mental workload associated with using technical GenAI tools like the SD WebUI, noting ArtistUI’s assistance in ideation and enhancing the creative process:
This tool could make it possible to generate a large number of variations quickly, then propose a large selection of images to the art director and decide which ones can be used in production. (P1 Q2)
Sample quotes evaluating the artefact from different vantage points are presented in Appendix A Table A1.
5.2.2. External evaluation.
The external evaluation of Generative.ConceptCrafter was conducted with six concept artists who had professional experience (inclusion criteria) and were recruited via LinkedIn and Instagram. This summative evaluation aimed to assess the artefact’s practical utility, usability, and creative support in real-world design tasks outside the case company. Participants EP4 and EP5 had previously applied GenAI tools in their creative workflows. Using a counterbalanced within-subjects design, each participant was tasked with generating two character poses under three conditions: using their routine tools and workflows, using OpenAI’s DALL-E,6 and using Generative. ConceptCrafter. The order of the conditions was varied to reduce learning effects. Each condition was limited to 15 minutes, during which we recorded completion times and participants’ perceptions of each tool.
This design enabled comparison of performance, usability, and creative outcomes across different tools while controlling for participant-level variation.
All participants utilised the full 15-minute allocation when working with their ordinary tools to create a character sheet with two poses, with one participant (EP6) unable to complete the second pose. Participants estimated they would require approximately one hour to complete the brief with proper presentation and colouring using conventional methods. However, all participants successfully completed the task under conditions two and three. On average, participants spent 10:55 minutes (σ = 3:13) using DALL-E and 11:30 minutes (σ = 3:47) using Generative. ConceptCrafter to achieve acceptable results.
Following task completion, we conducted semi-structured interviews to collect qualitative feedback on participants’ experiences, focusing on creative efficiency, output quality, and tool usability. Participants also reported their years of professional experience and proficiency level with GenAI tools. This complementary data provided a comprehensive understanding of how Generative.ConceptCrafter performs in real-world scenarios and yielded rich empirical insights from experienced concept artists for evaluating the tool’s design and practical utility outside the case company.
External participants acknowledged Generative. ConceptCrafter’s robust pose generation capabilities, which were lacking in DALL-E. The lack of control in off-the-shelf tools like DALL-E emerged as a shared concern among participants. The external evaluation demonstrated that both DALL-E and Generative. ConceptCrafter, as CSTs, can improve workflow efficiency in character generation tasks. Generative. ConceptCrafter provides artists with greater control through its customisable design capabilities, thereby enhancing the user experience for concept artists (see Appendix F Table A6).
Researchers and concept artists used output evaluation criteria such as relevance and accuracy (alignment with prompt), consistency (ability to produce similar outputs across similar prompts), visual quality (absence of defects or hallucinated elements), and creativity/originality (novelty and ideation potential) to assess image generation performance. While these metrics were not presented directly to users as in-system scores, they served as a basis for qualitative feedback and artefact refinement.
6. Design principles for text-to-image GenAI.
creativity support tools
6.1. Reflection and learning.
The ADR framework suggests that in parallel with the problem formulation and BIE stages, researchers should continuously engage in reflection and learning to move conceptually from building a specific solution to applying knowledge outcomes to a class of problems. The ADR principle “guided emergence” highlights that the artefact should resemble not only the preliminary design but also the continued influence of organisational use, perspectives, participants, and concurrent evaluation. As previous sections show, due to the reciprocal shaping nature of the organisational (the role, responsibilities, opinions, and routines of concept artists in game development pipelines) and technological forces (GenAI), the artefact changed and became more refined throughout the study.
In reflection and learning, we observed three helpful practices facilitating the artefact’s guided emergence and our reflections on the experiential process of building the artefact. First, because we deployed the artefact for routine concept art generation within and outside the case company with concept artists of varying experience levels, we were able to identify drawbacks and limitations (unanticipated outcomes) in an operational context. During routine tasks, concept artists could immediately benchmark the added benefits or hurdles introduced by the artefact in their creative endeavours.
Second, the ADR team engaged with the other members of the organisation by conducting meetings, presentations, demos, surveys, and interviews. The data collected from multiple channels and feedback obtained from participants with different perspectives (strategic, managerial, and operational) were crucial in capturing an unbiased, holistic view of the development and usage of GenAI applications to bring utility and value in organisational contexts. This also helped us to curate a diverse body of domain expertise in updating our artefact.
Third, the rigorous concurrent evaluation of the artefact (see Section 5) validated the usefulness of its features and the added value they bring to creative processes. As we reflected and learned about anticipated (e.g., enhancements in the concept art generation workflow) and unanticipated outcomes (drawbacks and limitations of the workflow and WebUI) that demanded ongoing changes to the preliminary artefact design, we identified a set of critical artefact features and design activities to enhance productivity of visual design processes through GenAI CSTs. These features and activities are process guidance, prompting, model parameter setup, output (image) performance metrics, and UI, and they are essential constituents of the design theory for text-to-image GenAI CSTs in visual design applications.
Based on our reflections and learning from the development and evaluation of the aforementioned integral aspects of the text-to-image GenAI CSTs, we formalise our learning as a set of DPs.
6.2. Formalisation of learning.
We formulated our DPs by following the DP template proposed by Chandra et al. (2015) and recently used by Seidel et al. (2018) and Pan et al. (2021). According to this template, a DP contains three types of information: material properties of the artefact, action potentials they offer, and boundary conditions. The template is as follows:
Provide the system with [material property—in terms of form and function] in order for users to [activity of users—in terms of action], given that [boundary conditions—user group’s characteristics or implementation settings].
Our DPs are as follows:
DP1 (process guidance): Provide users with an instructive generation workflow, including model parameters (e.g., sampling method, denoising, steps, batch size, scaling, etc.), model controls and guardrails, and prompt templates [material property], so that they can refer to it as guidance [action potential] when navigating the (technical) complexities of text-to -image GenAI CSTs for visual design [boundary condition].
DP2-a (prompting): Provide features to prompt the system with relevant and specific keywords arranged in keyword categories (e.g., subject, media, camera, topic, and extra tags) sequenced in the optimal order [material property], so that users can structure their prompts such that it increases the likelihood of generating the desired outputs [action potential] in text-to -image GenAI CSTs for visual design [boundary condition].
DP2-b (prompting): Provide the system with a negative prompt [material property], so that users can avoid undesired elements and retain more control over the generated output [action potential] of text-to-image GenAI CSTs for visual design [boundary condition].
DP2-c (prompting): Provide features that allow users to include reference images and apply structure-preserving extensions such as ControlNet [material property], so that users can condition the generation process on both visual and textual inputs to better communicate their creative intent and guide visual output generation [action potential] in text-to-image GenAI CSTs for visual design [boundary condition].
DP3 (model parameters): Provide features to choose the sampler and model checkpoints based on generation priorities of that workflow step [material property], so that users can decide on the speed-quality trade-off and the artistic direction [action potential] of text-to-image GenAI CSTs for visual design [boundary condition].
DP4 (output performance metrics): Provide users with output performance metrics such as accuracy, consistency, visual quality, creativity, and originality [material property], so that they can critically evaluate the performance (against human creativity as a benchmark) [action potential] of text-to-image GenAI CSTs for visual design [boundary condition].
DP5 (UI): Provide a streamlined UI that presents the prompt, negative prompt, sample image, and generated output, while handling model configurations in the backend [material property], so that users can convey their intent to GenAI models, conveniently bypassing intricate model configurations and avoiding long prompts [action potential] in text-to-image GenAI CSTs for visual design [boundary condition].
Table 5 presents the rationalisation of the DPs, their relation to our artefact, and the challenge addressed.
While each DP stands on its own, we recognise important relationships and interdependencies across the set. Several DPs operate at different stages of the creative process: DP1, DP2-a, DP2-b, DP2-c, DP3, and DP5 primarily support idea generation, focusing on usability and engagement, while DP4 is oriented towards idea selection, enabling critical evaluation of generated outputs. There are also functional synergies. e.g., DP1 (process guidance) is more effective when paired with DP5 (simplified UI), and DP2-a (structured prompting) complements DP2-b (negative prompting) and DP2-c (multimodal prompting) to produce more targeted and refined outputs. Together, these prompting principles enhance the usefulness of DP4 by enabling more meaningful output assessment.
Importantly, our DPs operate at the artefact feature or component level, guiding the design of interface elements, prompt engineering strategies, and parameter control mechanisms.
We adopted the Design Principle Reusability Evaluation Framework to assess the external validity and transferability of our DPs to other contexts within the class of text-to-image GenAI CSTs. The framework evaluates reusability across five key criteria: accessibility, importance, novelty and insightfulness, actability and guidance, and effectiveness. We engaged seven external practitioners—four CST developers and three human-AI collaboration and human-computer interaction (HAI/HCI) researchers—who reflect the target audience of our DPs, recruited via snowball sampling, a common method for recruiting experts in niche domains. As shown in Figure 9 and Appendix E, the findings indicate that the DPs are perceived as relevant, actionable, and valuable in guiding the design of GenAI CSTs beyond our specific case setting.
Q1 under the “importance” criterion, “in my view, Text-to-Image GenAI CSTs address real problems that often occur in my professional practice”, received a relatively lower score. We attribute this to a mismatch between the question and the professional backgrounds of the evaluators, who were CST developers and HCI/HAI researchers rather than concept artists.
7. Discussion.
The recent growth of GenAI has disrupted creative workflows in organisations. Having observed the superior performance of these generative models and the efficiency gains they bring, today’s organisations race to integrate GenAI applications into their creative processes. As we identified in this study, organisations and users face significant challenges when adopting general-purpose GenAI solutions in production-oriented professional settings, such as the skill gap, technical complexity, job insecurity and displacement, and errors and hallucinations. Despite the benefits reaped from GenAI implementations, these trade-offs fuel employee resistance, leading to a lack of user acceptance of GenAI systems and, ultimately, project failures.
GenAI being a novel phenomenon, we lack prescriptive design knowledge to navigate these trade-offs in developing and customising GenAI applications in creative industries like visual design. This lacuna in prescriptive knowledge not only begets costly failures of GenAI projects but also induces anxiety in practitioners when pursuing GenAI initiatives.
To address this gap, we develop a bridging artefact which assists concept artists to overcome creative blocks through rapid visual conceptualisation of diverse artistic ideas, ultimately improving their productivity, and we formalise seven DPs aimed at providing prescriptive guidance for integrating text-to-image GenAI CSTs into visual design workflows. To address the skill gap and the complexity of GenAI tools, we developed a process guidance workflow that makes the artefact accessible to users of varying expertise levels, instantiated through DP1. This ensures that both novice and experienced concept artists can effectively utilise the artefact, bridging the skill divide.
To mitigate issues related to resistance, deskilling, and concerns about job insecurity and displacement, our design emphasises effective human-GenAI collaboration through user-centred UI customisation and accessibility (DP5) and a well-crafted interaction design, particularly in terms of model parameter settings (DP3) and prompting techniques (DP2-a, -b, -c). These elements foster a collaborative environment where GenAI supports, rather than replaces, human creativity, ensuring that users feel empowered rather than threatened by the technology. Finally, the challenge of errors and hallucinations is addressed by incorporating structured prompting (DP2-a), negative prompting (DP2-b), multimodal prompting (DP2-c), structured parameterisation (DP3) and model performance metrics (DP4) to systematically evaluate generated outputs.
In this way, we ensure that the artefact generates reliable and accurate results, minimising the potential for erroneous outputs that could undermine user trust.
Through these DPs, we offer prescriptive design knowledge for integrating text-to-image GenAI CSTs into visual design workflows, including the development of key artefact components and customisation of off-the-shelf general-purpose tools like SD while overcoming the practical challenges organisations face. The design theory thus developed guides practitioners in making design choices in process guidance, prompting, model parameter setup, output (image) performance metrics, and UI, which are integral elements of text-to-image GenAI CSTs.
7.1. Implications for human-AI collaboration.
Our class of systems, text-to-image GenAI CSTs in visual design, and the DPs targeted at this class have several significant implications for human-AI collaboration scholarship and practice. First, the use of AI has sparked debates about overreliance on AI leading to human deskilling. Creativity, individual and collective, is a crucial human skill that is instrumental in enhancing the innovation capacity and competitiveness of firms. Before the recent advancements in GenAI, the scholarly discourse agreed that humans outperform AI on creative tasks and generativity. Hence, creative tasks in organisations were predominantly undertaken by humans, while prediction tasks were delegated to or augmented by AI. Lindebaum et al. (2020) found that such use of algorithmic decision making in organisations could hamper its overall experimentation and creativity capabilities.
Today, however, with GenAI models becoming increasingly powerful and accurate, it is questionable whether humans perform more effectively and efficiently than GenAI in creativity. With the rush to integrate GenAI into the creative processes of organisations and the number of GenAI applications available at the fingertips of employees, concerns about loss of creativity in humans have been raised. Jarvenpaa and Välikangas (2020) argue that what they call “inner time”—time to reflect on actions, meaning, and consequences—and “social time”—time to practise giving and taking multivocal ideas and perspectives—are vital for collaborative creativity. As we noted above, the increased use of GenAI might suppress both inner time, due to rapid generation cycles that do not allow sufficient time to reflect, and social time, due to engagement with
GenAI tools rather than humans in the creative processes. Citing these imminent threats of creativity deskilling, scholars of human-AI collaboration discourse have called for increased human involvement and engagement. The DPs of this study promote active human involvement and engagement by outlining material properties and the user actions they enable in GenAI artefacts.
Second, interactions between GenAI applications and users foster ample co-learning opportunities. When using GenAI applications, on the one hand, GenAI learns the user’s intentions through prompting and priming. Priming GenAI refers to setting an initial context or providing prior instructions to influence GenAI’s output by further contextualising the output. A preliminary stage provides the GenAI model with a direction, which guides its output and overall interaction. Through multimodal prompting, users can communicate precisely and intuitively their creative intent to the GenAI model (Z. Wang et al., 2024; Wu et al., 2024). On the other hand, the users can reflect on the generated outputs to learn from them and develop novel ideas. A lack of reflexivity of users hinders their and their organisations’ learning.
Mutual learning between GenAI and the user reinforces each agent’s strengths and buttresses the complementarities between human and GenAI. While our DP2-a, -b, and -c guide the learning of GenAI by delineating artefact features and user activities, we emphasise the importance of reflexivity, self-exploration and expression, experimentation, and critical engagement of the user.
Third, there are different types of creativity requirements in organisations. For instance, in our context of gaming, radical creativity is required to create a completely new character, but in other cases, incremental additions to the character or environment suffice. The generated outputs of GenAI applications may lack diversity and variability, precluding radical innovation and outside-the-box thinking. This inherent limitation stems from the fact that GenAI models struggle to generate outputs that are significantly different from their training data. Hard coding model parameters to simplify the UIs and hide workflow complexities further exacerbates the lack of outcome variability and hence radical creativity. In this study, we compromised between simplicity offered to the users and radical creativity.
While we endorse simplicity in UIs and concealment of complexities from the users (DP5), we accept that this approach is aimed at promoting incremental creativity and may curtail radical creativity/innovation.
Fourth, as design science research strives to provide innovative artefacts to solve practical problems, it is paramount that these artefacts achieve practitioner relevance and user acceptance. Algorithmic aversion is a prominent topic in the human-AI collaboration literature. There are multiple reasons for algorithmic aversion, such as task complexity, domain knowledge, AI experience, and cognitive biases. Humans tend to avoid using AI for complex and creative tasks compared to using it for mundane tasks, seeing it as a competitive threat, propelling their replacement with AI (G. Wang et al., 2024). Knowledge workers with higher experience also tend to exhibit greater aversion towards AI. Conversely, expertise in AI increases the willingness to engage with it. Recent design science research has responded to calls for increased practitioner relevance and user acceptance of AI artefacts.
Our DPs, especially DP1, 3, and 5, aim at filling practitioners’ needs and increasing user acceptance. Specifically, we lower usage barriers for users inexperienced in AI and art. Hence, novice artists can use our artefact with minimal AI training and produce art rapidly.
Fifth, the intention-to-prompt gap remains a significant challenge in contemporary text-to-image GenAI models (Z. Wang et al., 2024; Wu et al., 2024). Our bridging artefact addresses this by refining and optimising prompts provided by concept artists through structured prompt engineering. Specifically, we categorise and emphasise keywords using attention values, sequence them in an empirically validated order (DP2-a), and integrate negative prompts using presets and templates tailored for game character development (DP2-b). We extend this approach through multimodal prompt engineering (DP2-c), enabling users to condition outputs on reference images via ControlNet, capturing spatial cues. As crafting aesthetically compelling images often requires specialised domain knowledge, our process guidance workflow (DP1) equips users with essential GenAI and visual design knowledge.
Together, these mechanisms support users in translating artistic intent into effective prompts and, ultimately, into intent-aligned visual outputs—even when their initial textual articulation is incomplete.
Finally, in addition to academic and practical implications, our artefact and the broader adoption of GenAI raise noteworthy ethical, environmental, and legal concerns. First, from an ethical standpoint, there is a risk of biased outcomes generated by GenAI models, potentially leading to reputational damage for organisations, as illustrated by the recent Google Gemini case. Moreover, the displacement of human artists by automated processes poses ethical questions regarding the impact on the labour market and the livelihoods of creative professionals. Second, environmental concerns arise due to the heavy computing and storage requirements of GenAI systems, which can result in increased energy consumption and contribute to environmental degradation.
Finally, from a legal perspective, there are challenges related to the use of data for training GenAI models, particularly concerning copyright permissions and intellectual property rights. Addressing these ethical, environmental, and legal considerations is crucial to ensure the responsible and sustainable deployment of GenAI applications in creative contexts.
7.2. Limitations and future research.
While our study provides valuable insights into the integration of GenAI into visual design processes, it is important to acknowledge certain limitations that present opportunities for future research. First, our DPs were derived from a single case study within the gaming industry, focusing specifically on the use case of concept art creation. As we reviewed above, GenAI-enabled CST use cases are diverse in nature, including graphic design, product design, art creation, and fashion design. The use of GenAI in creativity support across different scenarios may yield markedly different results. Even within the gaming industry, GenAI can be used in various use cases, such as background score development, video snippets, interactive non-player character scripts, and game narration. Our narrow scope may limit the generalisability of our findings to other industries and applications of GenAI.
Future research could explore these DPs in diverse creative industries, such as filmmaking, cartoons, music, and digital art, to provide a more comprehensive understanding of their applicability across different contexts.
Second, while we built a mediating tool to bridge general-purpose GenAI models with the specific needs of organisational workflows, we acknowledge that this is only one approach. Alternative strategies, such as fine-tuning, reinforcement learning from human feedback, or retrieval-augmented generation, may offer complementary or more scalable ways to adapt foundation models for production use. Future work could explore how these approaches compare in terms of usability, acceptance, controllability, and cost across different creative domains.
Third, the simplified UI (DP5) of our artefact may result in a lack of control over image generation, posing a limitation for the effectiveness of our approach. While simplification of the UI enables users to swiftly adapt to the use case by reducing the entry barrier, it also introduces constraints on the benefits of the technology. Using (standard) sophisticated UIs to offer greater flexibility and customisation options could address this limitation and improve the usability of GenAI tools for experienced creative professionals.
Lastly, given the rapid advancement of GenAI technology, there is a risk that our findings may become outdated relatively quickly. A large number of different GenAI applications are mushrooming on a daily basis and the frontier of GenAI capabilities continues to move. The scope and capabilities of IS artefacts directly impact the DPs under consideration. In order to align with the latest developments, continuous monitoring of technological developments and updates to our research findings will be imperative to ensure their relevance and applicability in a fast-evolving landscape. Addressing these limitations through further research and refinement of our approach will contribute to the ongoing discourse on the responsible and effective integration of GenAI into creative practices.
8. Conclusion.
This paper presents the development of a text-to-image GenAI CST for visual design, positioned as a bridging artefact designed through a user-centred approach within an organisational context. We derive seven prescriptive DPs for text-to-image GenAI CSTs for visual design and offer guidance on harnessing the potential of GenAI to enable efficient creativity support, thereby contributing to the emerging field of human-AI collaboration. These principles serve as a guidepost—avant la lettre—for navigating the complexities of integrating GenAI into visual design processes, offering prescriptive design knowledge for responsible and effective collaboration between humans and GenAI.
Notes
1. The content created in this study has not been used in.
any of Ubisoft’s games or products.
2. the linked source.
release.
3. the linked source.
4. DXL is more recent, trained on higher resolution.
images, but it takes longer to generate images.
5. A neural network architecture designed to add spatial.
conditioning controls to the image generated.
6. the linked source.
Acknowledgements.
We thank Ubisoft for their collaboration, the Senior Editor and reviewers for their constructive feedback, and Esther Le Mair for excellent copyediting. This research was reviewed and approved by the Research Ethics Committee of the University of Lausanne.
Disclosure statement
No potential conflict of interest was reported by the author(s).
Funding
This work was supported by the Swiss National Science Foundation under Grant 197763 and 215542.
CRediT author statement
Savindu: Conceptualisation, Methodology, Writing, Project administration. Amirsiavosh: Software, Investigation, Validation, Writing, Funding acquisition. Yonah: Software, Validation. Yash: Supervision, Reviewing, Editing.