The 8-item Patient Health Questionnaire, or PHQ-8, is widely used to assess depressive symptom severity. Because it is brief and easy to administer, it is widely used across research, healthcare, education, and public-health settings.
Yet widespread use does not guarantee that a questionnaire performs equally well in every population. Individual questions may carry different meanings, prompt different response patterns, or relate differently to the construct being measured. When those possibilities have not been examined, group comparisons become harder to interpret.
Before comparing scores, researchers therefore need to ask a more fundamental question: Are we using the same ruler?
This post accompanies our published article, “Psychometric Evaluation of the 8-Item Patient Health Questionnaire Among Filipino American Emerging Adult Men and Women,” in the Journal of Psychopathology and Behavioral Assessment.
Why population-specific evidence matters
Filipino Americans are often included within broad Asian American categories in psychological research. These categories can be useful for describing population patterns, but they may also obscure meaningful differences among the communities they contain.
Asian American populations vary in migration histories, cultural contexts, family systems, socioeconomic experiences, and relationships with mental-health services. Evidence from a broad Asian American sample cannot automatically be assumed to apply equally to every subgroup within it.
This matters in psychometric research. Including members of a population in a study is not the same as demonstrating that the study’s measures operate appropriately within that population.
Our study addressed this gap by examining the psychometric properties of PHQ-8 scores among Filipino American emerging adults and testing whether the measure operated comparably across ethnicity and sex.
What does it mean to use the “same ruler”?
Imagine comparing two groups’ heights using rulers that look identical but may not use the same scale. Meaningful comparison is impossible unless their markings represent the same units.
The same principle applies to psychological measurement. A depression score is constructed from responses to questions intended to represent an underlying psychological attribute. Before comparing scores across groups, researchers need evidence that the relationship between those questions and the underlying construct remains sufficiently consistent.
This property is examined through measurement invariance. It evaluates whether a measure has a comparable structure across groups and whether its items contribute to the overall score in similar ways. Without that evidence, a higher score in one group could reflect greater depressive symptom severity, but it could also partly reflect differences in how particular questions function.
What we examined
We compared a one-factor model, in which all eight PHQ-8 items represented one overall dimension of depressive symptom severity, with two alternatives that separated symptoms into affective and somatic dimensions.
We then tested whether the retained structure was comparable across Filipino American and European American participants, men and women, and ethnicity-by-sex groups. Finally, we examined whether PHQ-8 scores related to depression, anxiety, perceived stress, brooding rumination, self-evaluation, and quality of life in theoretically expected ways.
What we found
The results supported retaining a one-factor structure for the PHQ-8.
Although both two-factor alternatives demonstrated acceptable overall fit, the affective and somatic factors were highly correlated across the groups examined. The proposed dimensions therefore did not provide a clearer interpretation than a single overall depression factor.
Measurement invariance analyses also supported comparable measurement properties across ethnicity, sex, and the ethnicity-by-sex comparisons included in the study.
In practical terms, the PHQ-8 items appeared to represent the same underlying construct across the groups examined. The results supported the level of invariance generally required for meaningful score comparisons. This strengthens the interpretation that group differences in PHQ-8 scores are more likely to reflect differences in depressive symptom severity than systematic differences in how the measure functions.
The external associations also followed expected patterns. Higher PHQ-8 scores were associated with greater depression, anxiety, perceived stress, brooding rumination, and self-deprecation, and with lower positive esteem and quality of life.
Together, the findings support interpretation of the PHQ-8 total score among Filipino American emerging adult college students represented in this sample.
What the findings do—and do not—establish
These results provide population-specific evidence, but they should not be interpreted as universal validation of the PHQ-8 for every Filipino American population or setting.
Participants were emerging adults from one university in Southern California, and most were born in the United States. The findings may not generalize automatically to older adults, adolescents, recent immigrants, non-college populations, or Filipino American communities in other regions.
Future research should examine the PHQ-8 across additional ages, migration histories, regions, socioeconomic contexts, and modes of administration. The present study provides a foundation, not the final word.
Making the analytical process visible
We also wanted the study’s analytical process to be open to inspection. Published articles necessarily condense complex analyses, which can make it difficult for readers to examine every modeling decision or reproduce the complete workflow.
We therefore made the de-identified analytic data and reproducibility materials publicly available. The repository includes the executable Quarto source, analysis code, generated tables and figures, a journal-formatted PDF supplement, and an interactive HTML version of the workflow.
These materials allow readers to inspect the analyses, reproduce the results, and adapt portions of the workflow for related psychometric research.
Published article:
Read the published article
Public data and reproducibility repository:
Explore the data and complete reproducible workflow
Interactive analytical supplement:
View the interactive analytical supplement
Representation also requires measurement
Improving representation in psychological research involves more than recruiting diverse samples. It also requires evaluating whether the instruments used to describe those samples function as intended.
Broad demographic categories can conceal unanswered questions about specific populations. Population-specific psychometric research helps make those questions visible and strengthens the interpretations researchers, clinicians, and institutions make from questionnaire scores.
Before comparing outcomes across communities, we should establish that the measurement itself supports the comparison.
We need to know that we are using the same ruler.