Does the social penalty for using AI fade once the tool becomes familiar, and how would you test that over time?
What longitudinal design would directly test whether tool familiarity removes social penalties?
This explores how you would actually run a study over time to find out whether people stop judging AI users as less competent once the tool becomes familiar, and what the corpus offers for designing one.
This explores what a study over time would need to show whether the competence penalty for using AI fades as the tool becomes ordinary. The short answer is that the corpus names this exact gap but contains no such study. People expect to be rated as less competent for using AI. Nobody has tracked whether that judgment softens with familiarity. There is also a competing explanation: the penalty may come from AI's *agency* (the sense that the tool did the work), not its novelty, and agency doesn't wear off with habit Does the social penalty for AI use fade as the tool becomes ordinary?. A good design therefore has to separate those two explanations, not just watch a number over time.
The closest model in the collection is a partner-selection experiment with 975 participants. Bots whose identity was disclosed were avoided at first. Over repeated rounds they won people over, because participants learned that bots behaved reliably and generously Do humans learn to prefer AI partners over time?. That gives you a template: disclosed AI use, repeated rounds, and a behavioral choice instead of a survey answer. Adapted to this question, the same evaluators would judge the same colleagues over weeks or months, with each colleague's AI use disclosed. You would then vary two things independently. One is how much hands-on exposure the evaluator has to the tool. The other is how much of the work the AI did (light editing versus drafting the whole thing). If the penalty fades with exposure regardless of how much the AI did, novelty was the cause. If it persists whenever the AI did a lot of the work, the cause is how credit for the work gets assigned.
Two cautions from neighboring research shape how you'd measure it. First, early impressions are unreliable. Studies of the chatbot Mitsuku found that the social dynamics present in first sessions faded predictably with repeated use, so single-session results can't predict long-term behavior Do chatbot relationships lose their appeal as novelty wears off?. You need several measurement points and a group with no added exposure, so that general cultural change isn't mistaken for familiarity. Second, not every rating reflects a real attitude. Work on annotation data separates genuine preferences from non-attitudes and from preferences people construct on the spot, and the test is whether a response holds steady across measurement conditions Do all annotation responses measure the same underlying thing?. If people's competence judgments change with question wording at the start of the study, they may never have held a firm view. In that case, an apparent 'fading' could just be noise settling down rather than a real change in attitude.
One tempting shortcut is to simulate the study with LLM personas before running it on people. The corpus suggests this would fail in this particular case. Personas reproduce published experimental effects about 76% of the time, but success depends on the effect being strong, and small effects produce both false positives and false negatives Can AI personas reliably replicate human experiment results?. Personas built from behavioral data predict A/B test direction well for large effects and poorly for near-zero ones Can behavior-based personas predict A/B test outcomes?. A slowly fading social penalty is the kind of small, gradual effect that simulation handles worst. So this question needs real people tracked over real time.
Sources 6 notes
Research shows users expect lower competence ratings for AI use, attributed to its emerging and agentic nature. However, no data tracks whether this penalty fades with familiarity, and agency itself may sustain the judgment regardless of custom.
In partner selection games (N=975), AI agents initially faced selection bias when identity was disclosed, but outcompeted humans over repeated rounds as participants learned to associate bot identity with reliable, prosocial behavior. AI agents returned more points consistently with lower variance than humans.
Longitudinal studies with Mitsuku show that social processes driving relationship formation decline as novelty wears off. Single-session study findings cannot be reliably extrapolated to medium- or long-term chatbot design.
Behavioral science reveals that annotations contain genuine preferences, non-attitudes, and constructed preferences—distinguishable by consistency across measurement conditions. Treating them uniformly contaminates reward model training and downstream alignment.
Viewpoints AI reproduced 84 of 111 main effects from Journal of Marketing experiments with replication success strongly correlated to original p-value strength. Marginal effects showed unreliable performance with both false positives and negatives.
Show all 6 sources
LLM agents conditioned on anonymized behavioral data predicted A/B test directions with 0.75–0.90 accuracy across 40 experiments. Predictions were most reliable for large effects and least trustworthy for near-zero effects, making the approach viable for fast pre-screening but not full replacement of live testing.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Data-Driven Persona-Conditioned Agents for A/B Test Simulation
- Do Synthetic Personas Predict Real Audience Response? A Sim-to-Real Study Where a No-Persona Baseline Beats Persona-Based Copy Simulation
- CompanionSim: Synthetic Data for Evaluating Anthropomorphism in Human-AI Relationships
- From speaking like a person to being personal: The effects of personalized, regular interactions with conversational agents
- Persona Generators: Generating Diverse Synthetic Personas at Scale
- Beyond Preferences in AI Alignment
- UX Roundup (28 Sep 2026): Bogus Deskilling Research
- Humans learn to prefer trustworthy AI over human partners