If productivity jumps after an AI rollout, how do you know the tool did it, not the calendar?
How can study design rule out calendar factors like seasonality or staffing changes?
This explores how researchers design studies so that a measured effect, such as a productivity gain after an AI tool rolls out, can't be explained by things that change over time, like busy seasons, new hires or shifting workloads. The corpus doesn't treat seasonality directly, so this answer draws on the design principles it does cover.
This explores how a study can show that an effect comes from the intervention and not from the calendar. One caveat first: none of these notes is about seasonality or staffing changes as such. What the collection does have is several notes on the same underlying problem, which is keeping the thing you're testing from getting mixed up with everything else that changes alongside it. The main lesson is that you can't fix a calendar confound in the analysis afterward. You have to design it out by making sure both conditions experience the same time period.
The clearest model is the live A/B test, which runs both versions at the same time and assigns people at random. Any holiday rush or staffing change hits both arms equally, so it cancels out. The note on behavior-based personas shows how much weight the field puts on this: LLM agents built from real user data predicted the direction of A/B test results 75 to 90 percent of the time. Even so, the authors present them only as a pre-screening tool, not a replacement for live testing Can behavior-based personas predict A/B test outcomes?. METR's developer studies use the same logic at a finer level. They randomize individual tasks to with-AI or without-AI, so the same developers in the same weeks produce both conditions. Their August 2025 study shows how that protection can break down without anyone noticing. Developers who didn't want to work without AI dropped out, and 30–50% held back the tasks they preferred to do with AI. Selection bias crept in through who took part and which tasks entered the study, even though the timing was well controlled Did developers opt out of METR's AI study because of selection bias?. Concurrent control removes calendar effects, but a randomized design can still be undermined by who chooses to take part.
A second idea is to vary one factor at a time. SchemeArena changes tool domains, goals, oversight and pressure independently across 400 scenarios, so any change in behavior can be traced to a single cause. Earlier studies changed several things at once and couldn't tell them apart Can independent scenario factors isolate what drives scheming?. Calendar confounds are the same problem in everyday form: a before/after comparison changes the intervention, the season and the team all at once. The persona-drift monitoring study is a useful example of testing timing itself as a variable. When it compared adaptive intervention timing against a fixed schedule, timing made no difference. All the benefit came from choosing *what* to correct Does monitoring help more by choosing what to correct than when to intervene?. If you suspect a time effect, you can build it into the design as something to measure instead of hoping it stays out of the way.
The third safeguard is to commit before you look. Spark-to-Paper requires researchers to state what evidence will count before any results come in Can separating judgment from verification improve research paper reliability?. The note on ad hoc prompt engineering shows the failure this prevents: when criteria shift during a study, they drift toward whatever the data happens to show Does iterative prompt engineering undermine scientific validity?. For calendar factors, that means fixing the comparison window, the baseline period and the outcome measures in advance. That way a slow August or a new team lead can't later serve as a convenient explanation for a result you didn't expect.
Taken together, ruling out the calendar is less about statistical correction and more about who gets what, and when. The strongest designs make time the same for both groups. For practical methods aimed specifically at seasonality, like staggered rollouts, interrupted time series or difference-in-differences, you'd need to look beyond this collection.
Sources 6 notes
LLM agents conditioned on anonymized behavioral data predicted A/B test directions with 0.75–0.90 accuracy across 40 experiments. Predictions were most reliable for large effects and least trustworthy for near-zero effects, making the approach viable for fast pre-screening but not full replacement of live testing.
METR's second developer trial produced unreliable speedup estimates because developers unwilling to work without AI opted out, and 30–50% withheld tasks they preferred to do with AI. This selection mechanism likely pushed estimates downward and missed high-uplift tasks, making the data "only very weak evidence" for the true productivity effect.
SchemeArena's 400-scenario benchmark varies tool domains, instrumental goals, oversight conditions, and pressure independently, enabling attribution of scheming behavior to specific factors rather than bundled changes. This factorization addresses a core limitation of earlier work that could not separate cause from effect.
Across 1,200 simulated conversations, behavior-specific monitoring reduced drift by 87%, while adaptive timing showed no advantage over fixed schedules. The monitor's value came from diagnosing which behaviors needed correction, not from deciding intervention timing.
Spark-to-Paper architects paper generation as composable skills that isolate model judgment from executable, verifiable operations and require evidence specification before results are observed, reducing dependence on model correctness for consistency.
Show all 6 sources
Iterative prompt revision by single researchers introduces individual bias, shifts evaluation criteria to match LLM capabilities rather than task requirements, and creates self-fulfilling feedback loops. A validated pipeline with inter-coder reliability and pre-specified criteria is required instead.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- LLM-REVal: Can We Trust LLM Reviewers Yet?
- From Prompt Engineering to Prompt Science With Human in the Loop
- Data-Driven Persona-Conditioned Agents for A/B Test Simulation
- Spark-to-Paper: End-to-End Research Paper Generation as a Composable Skill
- Prompting Against Persona Drift: Comparing Intervention Timing and Content in LLM-Simulated Conversations
- We are Changing our Developer Productivity Experiment Design
- SchemeArena: Factorized Stress Testing of Scheming in LLM Agents
- Agent A/B: Automated and Scalable A/B Testing on Live Websites with Interactive LLM Agents