Do full web pages beat markdown chat for LLM responses?
When LLMs generate complete interactive web pages instead of markdown text, do users prefer them? And how do they compare to pages built by human experts?
The paper reports a system where, instead of answering a prompt with markdown text, the LLM generates "a single fully-generated web page and a set of accompanying assets, such as images," rendered as-is in the browser. Against the standard markdown chat baseline, the authors find their generated pages "overwhelmingly preferred by humans," specifically "in 83% of evaluated cases." Measured against PAGEN, a new dataset of pages built by human expert teams for the same prompts, the system's pages are worse overall but "at least comparable in 50% of cases." The authors also report that this capability "is emergent, with substantial improvements from previous models" — older LLMs could not do this reliably, and the jump appeared with the newest generation of models.
The system has three parts: a server exposing tools (image generation, search) that the model can call; a large hand-written system prompt covering goals, planning guidelines, examples, and formatting/tooling instructions; and a set of lightweight post-processors that catch and fix common HTML/CSS/JavaScript errors after generation. The paper frames this as replacing a human product-manager/designer/engineer team with an "instant AI team" assembled per prompt, and argues the result is a shift from a "finite collections of texts" paradigm (fixed apps, fixed templates) to an "infinite catalog" of ephemeral, purpose-built interfaces.
This sits closest to Do generated interfaces outperform text-based chat for most tasks?, which reports a lower preference margin (70%+) for a narrower mechanism: LLM-selected interactive widgets built from a structured, finite-state representation with iterative generation-evaluation refinement. This paper's system instead generates an entire page end to end — HTML/CSS/JS, images, and all — stitched together by prompting and post-processing rather than a structured interface model, and reports a higher preference margin plus two findings the other note lacks: rough parity with human-expert output half the time, and the claim that the whole capability is new to the latest models rather than a steady improvement. It also contrasts with Do generated analysis UIs really work better than chat?, which found generated UIs add authoring rigidity and prompting overhead for analysts composing widgets iteratively; this paper's one-shot, fully generated page sidesteps that composition problem entirely, at the cost of being slow and allowing no user editing after the fact.
The excerpt does not say who the human raters were, how many comparisons the 83% and 50% figures rest on, or what the expert-built PAGEN pages were optimized for, so the comparisons' statistical weight is unclear. It also does not explain what "emergent" means operationally beyond pointing to Tables 3 and 4 the excerpt doesn't include, so the claim that older models simply cannot do this, rather than do it worse, is the authors' characterization rather than something the excerpt demonstrates directly. The paper itself names generation speed (a minute or two per page) and occasional JavaScript/CSS/HTML errors as open limitations, so the implication holds only for settings that can tolerate that latency and residual error rate.
Inquiring lines that read this note 5
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Does AI-assisted work increase total productivity or just shift time? Should GUI agents use structured screen representations instead of end-to-end vision? What prevents LLMs from applying their reasoning knowledge to improve outputs?Related concepts in this collection 3
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Do generated interfaces outperform text-based chat for most tasks?
Explores whether LLMs should create interactive UIs instead of text responses, and under what conditions users prefer dynamic interfaces to traditional conversational chat.
lower preference margin, different mechanism (structured widget selection vs. full page generation)
-
Do generated analysis UIs really work better than chat?
TaskArtisan investigates whether putting a GUI into LLM-assisted analysis workflows improves usability and clarity, and what trade-offs emerge when analysts need to modify or reuse generated interfaces.
contrasts one-shot full-page generation with the rigidity/authoring cost of composable widgets
-
Do generative UI tools actually implement their stated design rationales?
Explores whether generative UI tools build interfaces that match the design reasoning they provide. This matters because plausible-sounding rationales might persuade users to trust outputs without verification.
Qualifies: Design Theater finds generative UI tools fail roughly a quarter of stated design rationales, undercutting the quality implied by A's preference result
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Generative UI: LLMs are Effective UI Generators
- AI Meets the Classroom: When Does ChatGPT Harm Learning?
- The Decision to Verify: How Warmth and User Characteristics Shape Reliance on Conversational Agents for Information Search
- Experimental evidence of the effects of large language models versus web search on depth of learning
- Monitoring AI-Modified Content at Scale: A Case Study on the Impact of ChatGPT on AI Conference Peer Reviews
- What Makes a Good Natural Language Prompt?
- LLM or Human? Perceptions of Trust and Information Quality in Research Summaries
- Are Large Language Models a Threat to Digital Public Goods? Evidence from Activity on Stack Overflow
Original note title
generative UI full web-page responses are preferred over markdown chat output in 83 percent of cases and emerge only in the newest LLMs