Harnessing LLMs Without Surrendering Control: Delegation Boundaries in Visual Data Storytelling Authoring

Paper · arXiv 2609.25700 · Published September 22, 2026
Co-Writing and Collaboration

ABSTRACT Despite the emergence of large language models (LLMs) for visual data storytelling workflows, there are open questions about how authors decide what activities or tasks to entrust to them and what should be “protected” or maintained under human control. To investigate this, we interviewed a cohort of 12 expert visual data storytellers. Our analysis shows that participants rarely treated LLMs as autonomous storytellers. Instead, they tend to selectively delegate execution-oriented tasks to LLMs while retaining control over activities that shape narrative intent and story meaning. Our findings show that LLM assistance is most productive after human seeding and constraint-setting, and that it shifts labor from production to verification. We discuss design implications for boundary-aware authoring tools, data-grounded generation, low-fidelity ideation, and reporting practices for LLM-based visualization research. Supplemental materials for this paper are available at https://osf.io/hcnp6.

Introduction. LLMs and other forms of Generative AI are increasingly being woven into the practices of visualization researchers, designers, and authors [11,27]. In particular, visual data storytelling has emerged as a potential avenue for this type of work [8]. Prior work investigating the potential of human-GenAI collaboration for visual data storytellers has identified many potential activities and benefits. For example, Li et al. [13] interviewed 18 data workers (primarily researchers and business analysts) to explore their preferences for potential human-AI collaboration during storytelling planning, implementation, and communication, such as generating candidate story ideas and plot points, sourcing and summarizing background material and datasets, and producing and debugging code, and categorizing the potential roles an LLM can occupy during such collaboration (e.g., as a creator, optimizer, reviewer, or assistant). However, gaps remain for understanding current expert practice.

Discussion / Conclusion. Based on our analysis, we briefly summarize a set of design implications for creators, researchers, and tool builders navigating the shift toward LLM-assisted visual data storytelling: I1 Proactive boundary-aware authoring. Human-LLM storytelling is not a binary choice but often a nuanced negotiation, with delegation levels varying depending on several factors including the specific activities, perceived potential risks, and desired authorial control. Authors should proactively identify which story components should remain under human control to mitigate perceived risks, which adopts a position of preserving human agency without rejecting automation [3,9]. I2 Data-grounded generation and provenance-aware verification. Our participants treated model outputs as hypotheses to be checked rather than facts to be accepted, especially when outputs involved claims, causal explanations, or audience-facing content. This reframes verification as a part of authorship: maintaining ownership requires checking whether narrative statements and visualizations remain accurately grounded in data, context, and intent [19].

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

Why do some clarifying approaches produce understanding while others just satisfy? Why does polished presentation create unearned authority in AI outputs? Can local safety checks guarantee system-level behavioral safety? How can infrastructure records verify actual agent behavior? Does AI assistance promote real skill development or substitute for independent learning? What design and behavioral factors drive false consciousness attribution to AI? Do writers recognize when AI writing assistance alters their expressed stance? Why do people disclose to AI systems despite their artificial nature? What linguistic features distinguish AI-generated text from human writing most reliably? How should designers communicate what AI systems truly are and can do? When should work require human-AI partnership versus full automation? Why don't LLMs reliably translate capability into accurate outputs? What prevents conversational agents from taking initiative in dialogue? Do reasoning benchmarks predict model performance in long-horizon workflows? How does evaluation scope and dimensionality affect what we measure? How do standardized protocols improve multi-agent coordination and reliability? Should agents decouple planning from perception grounding for better performance? Can memory architectures handle ultra-long context better than attention?