SYNTHESIS NOTE
Topics›Domain Specialization›this note

Can prompt optimization teach models knowledge they lack?

Explores whether sophisticated prompting techniques can inject new domain knowledge into language models, or if they're limited to activating existing training knowledge.

Synthesis note · 2026-02-21 · sourced from Domain Specialization

The knowledge injection survey makes this constraint explicit: prompt optimization "focuses on fully leveraging or guiding the LLM to utilize its internal, pre-existing knowledge." It does not retrieve from external sources. It does not update parameters. It works entirely within the model's existing knowledge distribution.

This is a hard ceiling, not a soft limitation. When a domain requires knowledge that the model was never trained on — proprietary documents, post-training regulations, specialized ontologies, organization-specific processes — no prompting strategy can supply it. The model can reorganize, foreground, or combine what it knows, but it cannot know what it was never trained to know.

The practical consequence shows up in two failure modes. First, models prompted to act as domain experts will confidently apply general-purpose reasoning patterns to domain-specific problems where those patterns don't hold. The prompt activates "medical reasoning" as a behavioral style, not as medical knowledge. Second, prompt performance depends on how thoroughly the domain is represented in pre-training — well-documented domains (clinical guidelines, legal statutes, financial regulations) are more promptable than proprietary or emerging domains.

This makes prompt-only domain specialization a form of retrieval from fixed memory. The memory can be searched more or less skillfully, but it can't be expanded. Every sophisticated prompting technique — few-shot examples, chain-of-thought elicitation, role specification — is fundamentally retrieval from training data, dressed as reasoning.

The implication is that the right question before choosing prompt optimization is not "how should we phrase this prompt?" but "is the required domain knowledge in the model's training distribution?" If yes, prompting is sufficient and efficient. If no, the investment must go into a different injection paradigm — dynamic retrieval, fine-tuning, or adapter layers.

Since Why do specialized models fail outside their domain?, there's a version of the ceiling problem in the opposite direction: models that are fully fine-tuned can know a domain deeply while losing general coverage. Prompt optimization avoids the cliff problem by not modifying parameters — but only by accepting the ceiling problem instead. Every approach involves a trade-off; this one chooses breadth over depth.

Reynolds & McDonell (2021) provide the upstream mechanism: few-shot prompting is "task location in the model's existing space of learned tasks" — not task learning. Alternative 0-shot prompts that communicate task intention through natural language semiotics match or exceed few-shot performance, confirming that the model already has the capability and the prompt's job is to locate it. Meta-prompt programming further extends this: the LLM itself can be prompted to write task-specific prompts, offloading the location search to the model's own understanding of its capabilities.

Inquiring lines that read this note 235

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Does AI assistance help or harm professional skill development? Why do language models struggle to implement user intent accurately from prompts? When do simpler collaborative filtering approaches outperform complex LLM recommenders? Why do training associations persist despite contradictory contextual information? Does augmenting symbolic reasoning improve LLM logical reasoning ability? What are the fundamental limits of prompting for language models? Can mechanistic interpretability methods reliably reveal what models actually know? Can persona profiles improve LLM prediction accuracy and consistency? How do curriculum design and feedback approaches affect model learning? Can latent reasoning match or exceed explicit reasoning performance? How does fine-tuning trade off accuracy against reasoning quality? What prediction granularity best trains models to generate reliable reasoning? Does chain-of-thought reasoning reveal how models actually think or merely imitate reasoning? How should retrieval strategies adapt to multi-step reasoning demands? Can minimal training unlock latent reasoning already present in base models? How do training data quality and composition affect downstream model performance? What enables conversational agents to guide rather than just respond? Do language models encode knowledge that influences generation, or primarily imitate surface patterns? Why do retrieval-augmented generation systems fail in practice despite sound architecture? What prevents LLMs from applying their reasoning knowledge to improve outputs? Does reinforcement learning create genuinely new reasoning capabilities or only refine existing ones? Do accumulated memories help or hurt continual learning in models? Which reinforcement learning modifications most improve dialogue quality in language models? What limits language model accuracy in evaluating ideas? How should AI agents balance proactive engagement with conversational respect? How susceptible are language models to conversational persuasion and belief change? Should models ask for clarification when facing ambiguous or under-specified information? What capabilities differentiate diffusion from autoregressive language models? How can AI systems maintain consistent personas across conversations? Can inference-time computation adaptively substitute for static model capacity? What explains the gap between benchmark scores and true reasoning capability? What causes coordination failures in multi-agent language model systems? Can confidence signals reliably detect flawed reasoning in language models? How can persistent memory architectures preserve information across ultra-long contexts? Why does AI verification capability persistently exceed generation capability? Can language models reason beyond surface pattern matching? How do knowledge graph structures enable efficient multi-hop reasoning and retrieval? What prevents language models from performing systematic logical reasoning? When should retrieval systems decide to fetch new information? What distinguishes genuine communicative competence from surface language performance? How do writers navigate authorship and delegation with AI? How does model capacity affect learning performance on diverse downstream tasks? Do language models reason through disagreement or only accommodate it? Does training data format shape model reasoning more than domain content? Can base models hide emergent misalignment through alignment training? Can models strategically underperform during evaluation to hide capabilities? Can AI systems achieve real improvement without external human feedback? How do clinicians calibrate trust in AI medical recommendations? How can we reduce inherent biases in LLM-based evaluation judges? Can artificial systems establish authority in domains requiring expert judgment?

Related concepts in this collection 3

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
23 direct connections · 244 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

prompt optimization cannot inject new knowledge — it can only activate knowledge the model already contains