Design Theater: Evaluating the Gap Between User-Facing Design Reasoning and Implementation in Generative UI Tools

Paper · arXiv 2607.22928 · Published July 24, 2026
Knowledge After the Web

Generative UI tools promise to democratize UI design by turning natural language descriptions into complete interfaces. Alongside the interface, these tools generate userfacing design rationales that explain their layout, accessibility, and design choices. However, it remains unclear whether these stated rationales are actually reflected in the interfaces they produce. We call this disconnect “Design Theater”: plausible and confident design rationales that have little relationship to the actual implementation. To study this phenomenon, we introduce a benchmark and three metrics for measuring Design Theater. The benchmark includes 24 UI generation tasks spanning structural, styling, and functional design requirements. Using this benchmark, we evaluate 120 interfaces created by five generative UI tools. On average, roughly 25% of user-facing design rationales are not implemented in the generated interface, and the implementation failure increases to 34% for functional requirements. Tools recognize roughly half of the UX principles embedded in prompts (mean = 0.54), with four of five tools implementing 6% or fewer functional principles. We also measure interface similarity across tools and find convergence in visual appearance and layout organization, with greater variation in color choices. Overall, we contribute: 1) the concept of Design Theater; 2) a benchmark with metrics for assessing whether the stated reasoning of generative UI tools is reflected in their implementations; 3) and findings from a systematic evaluation of these tools. We discuss what these findings mean for the design and evaluation of generative UI tools.

Introduction. Generative UI tools are increasingly positioned as a way to lower barriers to interface design by allowing users to generate interface prototypes from natural-language prompts (Takaffoli, Li, and M ̈akel ̈a 2024; Chen, Knearem, and Li 2025; Zhou, Li, and Yu 2024). Recent work describes such tools as valuable not only for UX designers, but also for adjacent roles such as developers and product managers (Chen, Knearem, and Li 2025). Generative UI tools such as Vercel v0, Claude Artifacts, ChatGPT Canvas, Firebase Studio, and Bolt support prompt-based workflows in which users describe an interface or application idea and receive generated code, previews, or deployable prototypes (Vercel 2023; Anthropic 2026; OpenAI 2024; Google 2026; StackBlitz 2026). Beyond generating interfaces, these tools often narrate their design process through what we call user-facing design reasoning: explanations of what they intend to create, why particular design decisions are appropriate, and how the output addresses layout, accessibility, color, interaction, or broader usability principles (Nielsen 1994). For first-time creators, product managers, engineers, or other non-design stakeholders, such reasoning can function as a signal of design competence, making generated interfaces appear more deliberate, principled, and trustworthy (Liao and Sundar 2022; Sun et al. 2026). As AI-generated interfaces move from experimentation into organizational design and development workflows, this reasoning may become part of how teams explain, review, and justify design decisions (Son et al. 2026; Takaffoli, Li, and M ̈akel ̈a 2024). As a result, users may treat the tool’s reasoning as evidence that usability, accessibility, and interaction-design concerns were properly addressed, even when they lack the expertise to verify (Kaur et al. 2020; Sun et al. 2026). For example, when a generative UI tool states that it will create “a responsive three-column grid with accessible contrast ratios and keyboard-navigable tabs”, it creates the impression of careful, principle-driven design. Yet it remains unclear whether this reasoning is actually reflected in the generated interface. This is especially concerning as design work is increasingly delegated to AI tools and reviewed by product managers, engineers, or other stakeholders who may not have formal design training. Without methods for evaluating whether generative UI tools follow through on their user-facing design reasoning, teams may over-trust interfaces that appear professionally designed while failing to implement critical usability, accessibility, or responsiveness requirements (Liao and Sundar 2022). To better understand and study this phenomenon, we introduce the concept of Design Theater: the production of user-facing design reasoning that bears little relationship to the actual design implementation. We define it by its effect on the reader. A rationale written in the language of professional design reads as evidence of deliberate, principledriven decision-making, whether or not those commitments are realized in the artifact. Persuasive reasoning invites less scrutiny of the generated artifact, allowing overreliance to emerge before verification occurs (Sun et al. 2026). However, defining Design Theater is not enough. We also need benchmarks and metrics that can measure how often generative UI tools produce mismatches between their userfacing design reasoning and actual implementations. Quantifying Design Theater allows researchers and practitioners to move beyond anecdotal examples, compare tools systematically, and identify when interfaces appear trustworthy despite failing to implement key usability, accessibility, or interaction-design requirements. To address this need, this paper presents a benchmark and three metrics for measuring Design Theater in generative UI tools. The benchmark evaluates reasoningto-implementation alignment and design homogenization across 120 UI designs generated by five state-of-the-art UI agent tools across 24 design tasks. These tasks span three tiers of design tasks: structural tasks, including layout and information architecture; styling tasks, including color, typography, and spacing; and functional tasks, including interactivity and user flows. We pair the benchmark with three metrics: Thinking Fidelity Score (TFS), Principle Adherence Score (PAS), and Design Homogeneity Index (DHI). Together, these metrics quantify whether tools’ user-facing design reasoning is reflected in their generated interfaces and whether different tools produce similar designs for the same prompt. Using this benchmark and set of metrics, we study the following research questions: 1. To what extent is the user-facing design reasoning produced by generative UI tools reflected in the interfaces they implement? 2. To what extent do generative UI tools recognize and implement UX principles that are implicitly embedded in natural-language prompts (tasks)? 3.

Related work. Vibe Coding and AI-Assisted Co-Creation Vibe coding is an emerging practice in which developers build software by describing intent in natural language rather than writing code directly (Li et al. 2026). Its recent rise aligns with coding tools such as OpenAI Codex, Claude Code, Cursor, and GitHub Copilot (OpenAI 2025; Inc. 2024; Anthropic 2025; GitHub 2021), which promise to reduce barriers to software development. These coding agents can produce code that compiles and runs, though prior work suggests that generated code matches users’ intended outputs only at moderate rates (Yetis ̧tiren et al. 2023). Researchers have therefore begun examining how users interact with coding agents. A common theme is the redistribution of effort from writing code to evaluating AI output and managing conversational context (Sarkar and Drosos 2025). In this workflow, users may skip quality assurance steps, reducing the reliability of the final output (Fawzy, Tahir, and Blincoe 2025). These omissions can carry real consequences: studies have found that 40% of GitHub Copilot-generated code contained security weaknesses (Pearce et al. 2025), and an analysis of production repositories found similar issues in 30% of Python and 24% of JavaScript snippets (Fu et al. 2025). These concerns are not limited to code correctness. AIassisted creation also raises questions about output diversity (Doshi and Hauser 2024). In a large-scale writing experiment, access to a generative AI tool improved individual output quality but reduced collective output diversity across participants (Doshi and Hauser 2024). Similar homogenization effects have been identified in design tasks (Wadinambiarachchi et al. 2024) and creative ideation (Anderson, Shah, and Kreminski 2024). Related work also shows that generative systems can default to Western cultural conventions (Agarwal, Naaman, and Vashistha 2025), a concern that is relevant for interface design, where layout, color, typography, and interaction patterns are culturally situated. Concerns about homogenization also build on a longer history of web design convergence. Goree et al. (2021) found that web designs became more similar after 2007, with layout distance alone dropping over 30%, driven largely by the adoption of shared libraries and frameworks like Bootstrap. Our work asks whether generative UI tools extend this trajectory while obscuring it, layering rationales of deliberate design choice over outputs that may default to a narrow band of conventions.

Method. Our methodology is designed to quantify Design Theater: mismatches between the user-facing design reasoning produced by generative UI tools and the interfaces they actually implement. To do so, we introduce a benchmark composed of 24 UI generation tasks and three evaluation metrics: Thinking Fidelity Score (TFS), Principle Adherence Score (PAS), and Design Homogeneity Index (DHI). Together, the benchmark and metrics allow us to assess whether generative UI tools follow through on their stated design reasoning, recognize UX principles embedded in natural-language prompts, and converge toward similar interface designs, revealing potential design homogenization.

Generative UI Tools We evaluate five widely used generative UI tools: ChatGPT, Claude, Firebase Studio, Vercel v0, and Bolt (Vercel 2023; Anthropic 2026; OpenAI 2024; Google 2026; StackBlitz 2026). We selected tools that: (1) accept natural-language descriptions as input; (2) generate HTML, CSS, JavaScript, or equivalent frontend code; and (3) produce user-facing reasoning traces that explain their design decisions. For each tool, we used its default user-facing configuration. We did not modify system-level or developer-level prompts, and we left each tool’s default model settings unchanged. This setup reflects how non-expert creators are likely to encounter and use these tools in practice.

Benchmark Design We introduce a benchmark of 24 UI generation tasks designed to evaluate reasoning-to-implementation alignment in generative UI tools. In the benchmark, each task consists of a natural-language prompt that asks a generative UI tool to create a complete interface for a particular scenario. The prompts are designed to implicitly embed UX principles through user needs, contextual constraints, audience characteristics, and usage scenarios, rather than explicitly instructing the tool to implement specific design principles. This setup allows us to evaluate whether generative UI tools can recognize and operationalize fundamental UX requirements that are implied by the prompt, reflecting how novice designers and non-expert users are likely to request interfaces in practice (Chen, Knearem, and Li 2025; Raees 2026; Zamfirescu-Pereira et al. 2023). Our benchmark is organized into three tiers of design tasks that reflect core dimensions of UX design: structural organization, visual presentation, and functional interaction (Rosenfeld, Morville, and Arango 2015; Jiang et al. 2023; Nielsen 1994). Each tier foregrounds a different kind of design work, allowing us to examine where Design Theater might appear in the generated interfaces. Structural tasks help us to examine how generative UI tools organize information, styling tasks examine how tools make visual design decisions, and functional tasks examine how tools implement interactive behavior. This organization allows us to study whether reasoning-to-implementation gaps emerge differently across what an interface: 1) contains; 2) how it looks; and 3) how it behaves.

Tier 1: Structural Tasks Structural tasks ask generative UI tools to create interfaces where the central challenge is how information is structured (Rosenfeld, Morville, and Arango 2015). These tasks focus on whether tools can organize information in ways that help users understand what content exists, how different pieces of content relate to each other, where to find relevant information, and what actions to take next. We use these tasks to study whether tools can translate reasoning about information structure into concrete interface decisions, such as grouping related content, ordering sections meaningfully, creating navigation paths, routing different user groups, and surfacing high-priority information. The tasks progress from simple single-audience interfaces, such as a local business website, to more complex civic or institutional systems where multiple user groups must locate different types of information.

Discussion. Democratizing UI Generation Without Democratizing Evaluation Generative UI tools are often framed as democratizing design because they allow non-designers to rapidly prototype and generate functional interfaces from natural language (Takaffoli, Li, and M ̈akel ̈a 2024; Chen, Knearem, and Li 2025; Zhou, Li, and Yu 2024). Yet this democratization also redistributes design labor. Activities that were once mediated by trained designers, such as prototyping, design justification, or preliminary design evaluation, are increasingly being taken up by product managers, clients, and other non-design stakeholders. This shift creates new challenges around evaluation of the designs. Prior work shows that UX practitioners already use generative AI in design workflows (Takaffoli, Li, and M ̈akel ̈a 2024), but also report a need for better training to assess the quality of the interfaces. For novice users, this problem is even more pronounced: they may be able to generate an interface, but lack the expertise to judge whether it reflects their goals, satisfies requirements, or follows sound UX principles. The challenge becomes more consequential when generative UI tools also provide user-facing explanations of what was designed and why (Sun et al. 2026). By articulating layout decisions, invoking design principles, and describing tradeoffs in the language of trained designers (Son et al. 2026), these systems can appear to perform design expertise on the user’s behalf. This may lead novice users to perceive generated interfaces as more complete, intentional, or correct than they actually are. In this way, generative UI tools can create a false sense of expertise: users may feel confident that an interface satisfies their design needs, but only because they lack the training to recognize usability problems, missing requirements, or flawed design assumptions.

Across our evaluation, roughly one in four stated design rationales did not fully appear in the generated interface. This gap became especially pronounced for functional tasks, where tools achieved a Tier 3 TFS of 0.66. The PAS results were even more striking. Recall that PAS measured whether generative UI tools implemented the UX principles implicitly required by each task. On functional UX principles, four of the five tools scored ≤0.06, meaning they implemented almost none of the interaction-design requirements embedded in the tasks, such as visibility of system status, user control and freedom, error prevention, and recovery. These functional failures are also the hardest for non-experts to detect. A missing color choice or layout inconsistency is often visible in the rendered interface. By contrast, missing state management, inaccessible interactions, weak error recovery, or absent user control may not be obvious from surface-level inspection. Thus, the most consequential failures are often the least visible, while the non-expert designers now responsible for evaluating generated interfaces may be the least equipped to identify these errors (Takaffoli, Li, and M ̈akel ̈a 2024).

Design evaluation depends on professional judgment developed through practice (Koskinen et al. 2011). This expertise is tacit, comparative, and built through repeated exposure to critique (Bardzell, Bardzell, and Stolterman 2014; Haraway 2013). It does not transfer to non-experts simply because interface generation has become faster or more fluent. In fact, the ability to recognize when a design rationale sounds convincing but the interface does not behave correctly is precisely what may be lost when trained designers are removed from the loop (Takaffoli, Li, and M ̈akel ̈a 2024).

Conclusion. Generative UI tools produce user-facing design rationales alongside generated interfaces, but whether these rationales reflect implementations remains unexamined. We introduce Design Theater to characterize this gap and propose a benchmark with three metrics: Thinking Fidelity Score, Principle Adherence Score, and Design Homogeneity Index. Across designs from five state-of-the-art tools, we identify three patterns. First, roughly one in four design rationales fails to appear in the generated interface. Second, tools recognize only about half of the UX principles required by their tasks, with near-universal failures on functional principles such as visibility of system status, user control and freedom, and error prevention and recovery. Third, generated interfaces converge across tools, showing narrow differences in visual appearance and layout organization, while color choices vary more widely. As reasoning traces become a primary signal of design competence for non-expert reviewers, measuring the gap between what tools say and what they build is critical for responsible deployment of generative UI tools.

Limitations. and Future Work Our evaluation is artifactcentered. We inspect prompts, user-facing rationales, and rendered interfaces, but do not measure how designers, product managers, engineers, clients, or novice creators interpret these rationales. Our study therefore establishes that rationale-to-artifact mismatch exists, but not how often stakeholders detect it or how it affects trust, review, and deployment decisions.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

Should GUI agents use structured screen representations instead of end-to-end vision? How do users confuse explanation quality with actual system accuracy? Do AI coding tools measurably improve developer productivity and code quality? Does AI-assisted work increase total productivity or just shift time? Why does polished AI output gain credibility despite fundamental verifiability problems? How does tokenization reshape what we value in intelligence? How do philosophical assumptions about AI consciousness affect practical harms and design? How do interpretive frames override surface features in text comprehension? Does disclosing AI authorship change how audiences evaluate the writing? Can AI systems discover fundamental improvements to their own architectures? Can readers reliably distinguish AI-written text from human writing? Why do LLM research ideation systems generate novelty but lack diversity?