What came back were three different accounts of the same seven-day period, with women reporting roughly six hours of sitting a day against the 11.4 hours the accelerometer recorded. What a woman wrote on the questionnaire did not portray the same picture what her device recorded, so the participant who reported the most activity on the questionnaire was no more likely than any other to rank highest on the accelerometer. On a scale where one means perfect agreement and zero means no agreement at all, the self-reported values and accelerometer-measured values lined up at .29 for sedentary time, .03 for moderate activity, and .22 for vigorous activity. By six months even those weak associations had disappeared.
Those disagreements ran in a consistent and familiar direction, since the questionnaire overstated moderate to vigorous activity and understated sitting time at both visits. Other studies of breast cancer survivors have found the same pattern at larger magnitudes, with overestimates ranging from 147% to 994%.
The overreporting of moderate to vigorous activity grew from 86% at baseline to 118% at six months, after the women had spent half a year in a program encouraging them to move more.
“This study provided further evidence that either the way we think about how much physical activity we do, or how physical activity patterns are asked about in surveys, is not always accurate, especially for higher intensity exercise,” said study co-author Dr. Blake Langley, a staff scientist in the Public Health Sciences Division. “We tend to think we have higher moderate and vigorous physical activity than what is measured by devices we wear.”
The authors offer several explanations. The questionnaire asks only about particular domains such as work, transportation and housework, all of which the pandemic reshaped mid-trial. The other possibility is what’s called social desirability. According to this concept, the authors considered that since participants know that more activity and less sitting is the healthier answer, and after months of coaching they know it better than at the start, they may have over-reported their level of physical activity. These data cannot settle whether the women perceived more activity or reported what they had learned to say.
“We suggested that future studies that are reliant on self-reported measures, or that are looking to validate the accuracy or best approach in a specific group of people, should consider evaluating social desirability as that may be a reason study participants overestimate levels of physical activity,” Langley said.
Whether a woman counted as active enough depended on which measurement was considered. The accelerometer classified 40% of women as meeting the guideline of at least 150 minutes per week of moderate to vigorous activity at baseline while the questionnaire classified 59%. Among the 30 women with all three measures at six months, the figures were 50% by accelerometer, 73% by questionnaire and 67% by Fitbit. A tool that overestimates activity will quietly screen out the women who would benefit most from support.
For light activity the Fitbit correlated with the research accelerometer at .80, suggesting either device can capture the low-intensity movement that fills most of a day. Moderate activity showed no significant relationship, and the Fitbit put average vigorous activity at 25.7 minutes per day against the accelerometer's 1.2 minutes. Its intensity classifications come from proprietary algorithms that are not shared with researchers and change over time. Sedentary time could not be compared at all, because there was no way to tell when the device was off the wrist.
Accelerometers carry costs of their own. They are expensive to buy, mail and replace, and they demand staff time and expertise to clean and score the output. Seven days on the hip is a demanding ask that captures a single week, and there is still no standardized approach to processing the data.
“When we consider the surveys and questionnaires to include in our studies, we prioritize the experience for participants,” Langley said. “We don’t want to overburden participants with questions that don’t get to the root concern.”
“The solution to the measurement error isn’t straightforward,” he added. “Researchers will continue to balance time and cost burdens while ensuring the tools they select provide the best estimate for their purpose, including if a wearable device like a Fitbit watch can be used as a support tool for people looking to increase their physical activity.”
The takeaway is not that questionnaires should be abandoned. They cost little, run fast and ask almost nothing of participants, which is why surveillance studies depend on them. It is that the choice of instrument is a design decision with important consequences, and the instrument most likely to show an intervention working may not necessarily be the most appropriate measure for evaluating its effectiveness.