I’m not sure if this is what you were alluding to there, but I think there is a real problem with experimental research in this kind of field that often the test subjects are students, who tend to have much less skill and experience than working programmers with a few years of practice behind them.
It’s a specific instance of a more general problem: it seems clear that programmers often work differently depending on their familiarity both with programming generally and with the specific domain they’re working in, but we’ve only scratched the surface in identifying exactly how they work differently, and therefore what practical steps we might take to make things better for programmers in different situations.
This is often the problem with a lot of psychology research, where the only easily available subjects are university students who receive credit for participating in the research (or a nominal payment).