INQUIRING LINE

When an AI watchdog checks 167 signals for errors, which ones actually do the catching, and is the model's own confidence one?

Which of the 167 numerical features carry the strongest detection signal?

This explores which features in a detector built on 167 numerical inputs do the most work. The retrieved notes don't contain that feature set, so this reads the question more broadly: when a system tries to detect a problem (hallucination, reward hacking, a bad reasoning trace, AI-made content), which kinds of signals tend to carry the most weight?


This explores which features in a detector built on 167 numerical inputs carry the most signal. The direct answer is that none of the retrieved notes describes that feature set, so this synthesis can't rank those 167 features. It can say what the corpus has learned about which kinds of signals tend to work, and that can help you judge a list like this when you find one.

The first lesson is that features drawn from outside the model often beat the model's own confidence. Hallucinations about rare entities are easier to catch by looking at the pretraining data, such as whether two entities ever appeared together, than by asking how sure the model feels. Models are often most confident exactly when they are wrong about combinations they never saw Can pretraining data statistics detect hallucinations better than model confidence?. A related finding is that just 27 cheap features of the question itself can predict when a system should look something up. They match costly uncertainty methods overall and do better on hard questions Can question features alone predict when to retrieve?.

The second lesson is that the strongest single feature is often weaker than a pair of features that catch different kinds of failure. Model confidence and data rarity miss different errors: confidence misses mistakes about rare things, and rarity misses shaky reasoning about common knowledge. A trigger that combines the two beats either one alone Should RAG systems use model confidence or data rarity to trigger retrieval?. When and where a signal is measured matters too. Confidence measured at each reasoning step catches breakdowns that an average over the whole answer hides Does step-level confidence outperform global averaging for trace filtering?. So if a 167-feature study reports a single winner, it's worth asking whether that winner is really a stand-in for a combination, or a feature measured at the right level of detail.

The third lesson is that cheap signals can be surprisingly strong. A simple difference-of-means vector, read from computations the model already performs, catches about as many reward hacks as a separate LLM monitor at almost no extra cost How do cheap vector detectors compare to expensive LLM monitors?. Two warnings go with this. A model can contain every feature a task needs, readable with a simple linear probe, and still have a messy internal structure that breaks when inputs shift Can models be smart without organized internal structure?. And removing a feature that looks spurious can sometimes hurt performance rather than help Why does removing spurious cues sometimes hurt model performance?. Feature-importance rankings describe one dataset, not a law.

If the 167 features are meant to detect AI-generated content, here is why that work matters: people perform at roughly chance when trying to tell AI-made text, images, and voices from human-made ones Can people reliably spot content made by AI?. Measured features may be the only reliable signal available. To get a direct answer about the 167 features themselves, the source paper would need to be added to the collection.


Sources 8 notes

Can pretraining data statistics detect hallucinations better than model confidence?

QuCo-RAG uses entity co-occurrence patterns from training data to trigger retrieval, successfully flagging hallucination risk even when models are highly confident. This data-side approach catches the root cause (unseen combinations) rather than the symptom (low confidence).

Can question features alone predict when to retrieve?

Learned predictors using 27 lightweight external question features match complex uncertainty-based methods on overall performance while costing far less, and outperform them on complex questions across 6 QA datasets.

Should RAG systems use model confidence or data rarity to trigger retrieval?

Model confidence and data-rarity signals catch orthogonal failure modes: confidence misses hallucinations about rare entities, while rarity misses uncertain reasoning about common knowledge. Hybrid triggers substantially outperform either signal alone.

Does step-level confidence outperform global averaging for trace filtering?

Local step-level confidence catches reasoning breakdowns that global averaging masks and enables early stopping before traces complete. This approach achieves comparable accuracy gains to naive majority voting with far fewer generated traces, proving trace quality matters more than quantity.

How do cheap vector detectors compare to expensive LLM monitors?

On DeepSWE, difference-of-means vectors caught 3.1% more hacks in Kimi K3 but 7.9% fewer in GLM 5.2 than LLM monitors at matched false positive rates. The method applies to existing forward passes, making it virtually free compared to running a separate monitor model.

Show all 8 sources
Can models be smart without organized internal structure?

Models trained with SGD can contain all the linearly decodable features needed for a task while maintaining fundamentally broken internal organization. This makes them vulnerable to perturbation and distribution shift invisible to standard evaluation metrics.

Why does removing spurious cues sometimes hurt model performance?

Removing spurious cues degrades performance in heuristic override tasks, opposite to shortcut learning predictions. The failure mode is integrating conflicting signals rather than ignoring distractors—a frame problem, not feature selection.

Can people reliably spot content made by AI?

A 30-study systematic review found that humans cannot reliably distinguish AI-generated from human-created content across text, image, and voice modalities. Accuracy generally clusters around chance and has not kept pace with improvements in AI realism.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.