A note gets noticed by what it resembles, how it's filed, or the trail of clicks that lead to it — which matters most?
How do mimetic, administrative, and stigmergic adjacency shape visibility?
This explores how something becomes visible (to readers, search systems or AI models) depending on what it sits next to: things that look like other things (mimetic), things filed into the same official categories (administrative), and things linked by the traces other people's activity leaves behind (stigmergic).
This explores how being near the right things makes content visible: imitating what is already prominent, being sorted into an official category, or getting picked up by trails of other people's clicks and links. The retrieved notes don't answer this. They cover the internals of AI models, not the sociology of how attention moves through an information system. A real answer would need the material on knowledge after the web and platform visibility, and none of it came back for this question. What did come back shows the same three dynamics at work inside the machines, at a smaller scale.
Mimetic adjacency has the clearest parallel. Transformer attention gives extra weight to content that is repeated or already prominent in the context, whether or not it is relevant. That creates a feedback loop: whatever echoes what is already there gets amplified, before any training for helpfulness takes effect Does transformer attention architecture inherently favor repeated content?. In the models, then, resembling the crowd is itself a route to being seen. That is a mechanical version of the claim that imitation creates visibility.
Administrative adjacency shows up as structure imposed from outside. AI agents that operate computer screens do much better when the screen is first converted into a labeled list of elements, an official inventory of what is there, than when they look at raw pixels Can structured interfaces help language models control GUIs better? Why do vision-only GUI agents struggle with screen interpretation?. What isn't in the inventory is effectively invisible to the agent. Embedding spaces do something similar on their own: their main axes sort concepts into a top-down taxonomy, broad categories first and then finer ones, much like a dictionary's hierarchy Do embedding eigenvectors organize taxonomy from coarse to fine?. Which category something gets filed under decides what it ends up near.
Stigmergic adjacency, where visibility builds up from the traces others leave (citations, links, usage), has no counterpart in this retrieval. That gap is the most useful thing to take away. The corpus can tell you how models privilege repetition and categories, but these notes can't say how collective use leaves paths that make some knowledge findable and other knowledge invisible. If that is the thread you care about, ask about it directly in terms of search, recommendation or citation dynamics. It will retrieve better than this abstract three-way framing.
Sources 4 notes
Transformer soft attention systematically over-weights repeated and context-prominent tokens regardless of relevance, creating a positive feedback loop that amplifies opinions and framing before RLHF acts. System 2 Attention—regenerating context to remove irrelevant material—can interrupt this mechanism.
Agent S's dual-input design—visual input for environmental understanding plus image-augmented accessibility trees for grounding—achieved 9.37% improvement over baseline by factoring planning and grounding into separate optimization paths rather than forcing end-to-end prediction.
OmniParser demonstrates that GPT-4V fails when forced to simultaneously identify icon meanings and predict actions from raw screenshots. Pre-parsing screenshots into structured semantic elements with descriptions lets the model focus solely on action prediction, removing the composite-task bottleneck.
Leading eigenvectors of embedding Gram matrices separate broad taxonomic branches first, then progressively finer sub-branches—a coarse-to-fine spectral order that tracks the WordNet hypernym tree level by level, confirming predictions from co-occurrence statistics.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- ShowUI: One Vision-Language-Action Model for GUI Visual Agent
- OmniParser for Pure Vision Based GUI Agent
- MOMENTS: A Comprehensive Multimodal Benchmark for Theory of Mind
- Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents
- Hierarchical Concept Geometry in Language Models Emerges from Word Co-occurrence
- Agent S: An Open Agentic Framework that Uses Computers Like a Human
- BTL-UI: Blink-Think-Link Reasoning Model for GUI Agent
- ScreenAI: A Vision-Language Model for UI and Infographics Understanding