Table of Contents
- Key Highlights
- Introduction
- Study design and the surveillance dataset
- Between-class clustering: how large is the class effect?
- Individual-level correlates of composite fitness
- Class composition: compositional versus contextual effects
- Composite score versus BMI-defined abnormality: diverging surveillance signals (2020–2024)
- Why composite scores can mask weight trends
- Methodological considerations, diagnostics and robustness checks
- Practical implications for universities and health practitioners
- Recommendations for researchers and practitioners
- Future research directions
- The immediate takeaway for institutional surveillance
- FAQ
Key Highlights
- Analysis of 106,461 annual surveillance records from a Sichuan public university shows significant between-class clustering in composite physical-fitness scores (ICC ≈ 0.18), with both individual BMI and class-average BMI negatively associated with fitness.
- Composite fitness scores and the prevalence of BMI-defined abnormality moved in different directions between 2020 and 2024: composite scores rose and fell non-monotonically while BMI-defined abnormality increased steadily from 23.3% to 31.4%, demonstrating that a single weighted score can mask worsening anthropometric trends.
Introduction
Universities routinely measure student physical fitness, producing composite scores intended to summarise performance across endurance, strength, speed, flexibility and body composition. Those composite scores inform health-promotion decisions, resource allocation in physical education and institutional assessments of student wellbeing. Yet composite metrics can obscure important variation. When students are grouped into classes that share instructors, schedules, facilities and social norms, outcomes may cluster; when body mass index (BMI) contributes only a fraction of a weighted total, population-level shifts in weight status may escape notice even as the composite score appears stable or improving.
An analysis of five rounds of annual fitness surveillance (2020–2024) from one public university in Sichuan Province, China, illustrates both of these points. Using multilevel models on 106,461 de-identified records nested within 2,591 class-year units, investigators quantified how much variation in composite fitness sits between classes, how individual and class compositional factors relate to fitness, and whether changes in the composite score moved in step with changes in BMI-defined abnormality. The results show sizable class-level clustering, robust negative BMI–fitness associations at both individual and class levels, and a clear divergence between the composite score and BMI distribution across the surveillance window. The implications affect how universities analyse fitness data, structure surveillance reports, and interpret year-to-year trends.
The following sections describe the dataset and methods, present the core findings in context, examine methodological caveats, and outline what institutions and researchers should change in light of these patterns. Practical examples and recommended analytic practices are included to aid translation into policy and research.
Study design and the surveillance dataset
The source for this analysis was an institutional routine surveillance system operating under China’s National Student Physical Fitness Standards (NSPFS). Five autumn rounds—2020 through 2024—used the same NSPFS scoring framework and sex- and grade-specific norms. The NSPFS composite score ranges from 0 to 120 and allocates weight as follows: BMI 15%, vital capacity 15%, 50-m sprint 20%, sit-and-reach 10%, standing long jump 10% and an endurance/strength block 30% (gender-specific endurance tests). Testing occurred at the same university each year; identical scoring algorithms were applied across rounds.
Records included sex, date of birth, height, weight, component raw results, standardized component scores, and the total composite score. After screening for implausible BMI (outside 12–50 kg/m2), extreme ages (outside 15–35 years), unassignable grade, and very small class-year units (<5 records), the analytic sample contained 106,461 records from 2,591 class-year units. The median class-year size was 41 students; the pooled mean age was 20.7 years and mean BMI 21.6 kg/m2. Female records were more numerous (58.4%) and had higher adjusted composite scores than male records, though male students had higher average BMI.
Three class-level aggregates were computed within each class-year unit: class size, proportion male, and mean class BMI. Individual BMI was centered around the class mean to separate within-class (individual deviation) from between-class (class mean) associations. Calendar year entered models as a categorical fixed effect to accommodate non-monotonic patterns driven by pandemic-period disruption and return-to-campus effects.
Multilevel linear mixed models were fitted by maximum likelihood. A null random-intercept model estimated the intraclass correlation coefficient (ICC), quantifying the share of total variance attributable to class-year differences. Successive models added individual- and class-level covariates and then allowed the within-class BMI slope to vary across class-year units (random slope).
Complementary temporal analyses compared adjusted composite scores with standardized prevalence of BMI-defined abnormality (any category other than normal weight) for each year and assessed whether the two surveillance signals moved concordantly between adjacent years.
Between-class clustering: how large is the class effect?
The null random-intercept model returned an ICC of 0.178, meaning approximately 18% of variance in composite fitness scores resided between class-year units. Put differently, nearly one-fifth of the observed differences in composite scores reflect which administrative class and survey year a record belonged to, rather than only individual characteristics. The ICC remained meaningful in separate single-year null models (range ≈ 0.10–0.25), confirming that clustering was not an artefact of pooling across years.
Why does this matter? A substantial ICC implies that students sharing the same class environment produce correlated outcomes. That correlation may arise from shared schedules, the same physical-education instructors, common practice or training opportunities, peer exercise norms, or selection into classes. It might also reflect local administrative or testing conditions. Whatever the mechanism, ignoring clustering causes standard errors to be underestimated in single-level regressions and risks overconfident inferences about individual-level predictors.
A practical illustration: consider two class-year units in the same major. If one class consistently scores 8–10 points higher on the composite than another, that difference matters for resource targeting. Using a single pooled institutional average could mask such variation and misdirect interventions. A class with persistently lower scores might benefit from targeted instruction, facility access or tailored motivation programs even if the institutional average looks acceptable.
Individual-level correlates of composite fitness
Multilevel models that adjusted for class composition and calendar year identified several robust individual-level associations.
-
Within-class BMI: Students whose BMI exceeded their class mean scored lower on the composite. The within-class coefficient was approximately −0.53 points per BMI unit (kg/m2). This negative association reproduces extensive prior evidence linking higher BMI—especially overweight and obesity—to poorer performance on endurance, speed, and strength tests. Because BMI contributes 15% to the composite, part of the negative association is structural; the remainder reflects poorer component performance associated with higher adiposity.
-
Sex: Male records were associated with markedly lower composite scores after adjustment (approximate difference −4.17 points). Interpreting this coefficient requires caution. The NSPFS uses sex-specific scoring norms and different endurance/strength tests by sex, so the composite score mixes biological differences, test construction and mean performance. Year-to-year variation in the male–female score gap was substantial, pointing to both compositional and administration contributors.
-
Age and grade: Older age was weakly associated with lower composite scores. Cross-sectional grade contrasts suggested higher adjusted scores for second- and fourth-year records compared with first-year records, but those contrasts are between-pool comparisons, not within-student developments. Without stable cross-year identifiers the data cannot distinguish maturation, cohort selection, or curriculum effects.
Standardized coefficients ranked sex and within-class BMI as the strongest individual-level correlates. The fixed effects collectively explained about 15% of variance (marginal R2); the full multilevel model explained roughly 28% (conditional R2), leaving substantial unexplained variation potentially tied to behavioral, psychosocial or measurement factors not present in the surveillance files.
Real-world implication: Health promotion programs focusing solely on individual-level BMI may miss larger compositional dynamics. A student with above-class-average BMI is at greater risk of lower composite performance, but interventions will be more effective when they consider peer and class contexts.
Class composition: compositional versus contextual effects
Class-level aggregates—mean class BMI and proportion male—exhibited independent associations with composite fitness after adjusting for individual-level covariates and calendar year.
-
Mean class BMI: Units with higher average BMI showed lower mean composite scores (approximate coefficient −0.91 per unit of mean BMI). This association persists after accounting for a student's own BMI deviation from the class mean, highlighting that class composition matters in addition to individual status. Whether that association reflects composition (who is in the class) or contextual processes (shared behaviors, norms, or testing practices) cannot be resolved from aggregate data alone.
-
Proportion male: Classes with a higher proportion of male students had higher composite scores on average. That association is likely compositional—reflecting the scoring structure, gendered performance differences and how sex-specific components aggregate—rather than an environmental advantage conferred by more males per se.
-
Class size: Although statistically significant, the association between class size and composite score was practically negligible (approx. −0.012 per additional student). Across realistic class-size differences this translates into small expected score changes.
These findings reveal two practical consequences. First, class-level screening that relies on class-average composites should interpret low performance in light of compositional features. A class with many high-BMI students will, on average, score lower; an intervention targeting the class might therefore need to address both individual-level support and broader opportunities for physical activity and behavioral change. Second, reporting should separately present class-compositional statistics—mean BMI, proportion under/overweight—rather than only a single composite.
Illustrative example: two classes with identical average composite scores could differ sharply in composition—one might have many normal-weight but low-endurance students, another might have mixed BMI but strong sprinters who raise the speed component. Tailored interventions require disaggregated reporting.
Composite score versus BMI-defined abnormality: diverging surveillance signals (2020–2024)
A central result of the analysis is the lack of directional concordance between the NSPFS composite score and BMI-defined abnormality across the five-year window.
Composite score trajectory:
- 2020: mean 70.72
- 2021: increased to 73.37
- 2022: declined to 69.86
- 2023: rose to 71.00
- 2024: rose to 73.11
BMI-defined abnormality (underweight, overweight, or obesity) trajectory:
- 2020: 23.32%
- 2021: 24.15%
- 2022: 25.43%
- 2023: 28.81%
- 2024: 31.40%
The composite score moved non-monotonically, influenced by component-level changes and possible year-specific administration effects. BMI-defined abnormality rose steadily each year, culminating in a near 8-percentage-point increase from 2020 to 2024. Adjacent-year comparisons showed discordance in 2020–2021 (composite up, abnormality up), concordance in 2021–2022 (both worse), and discordance again in later intervals when composite scores recovered but BMI-defined abnormality continued to worsen.
The rising abnormality reflected both tails of the BMI distribution: overweight-plus-obesity grew from 16.9% to 23.3% while underweight, stable near 6% through 2022, rose to 8.1% in 2024. The increasing underweight fraction alongside growing overweight highlights the coexistence of divergent nutritional and behavioral risks within the university population.
Why does this matter? Surveillance that relies on a single weighted composite can produce an impression of stable or improving fitness while the BMI distribution shifts unfavorably. Decision-makers who view only the composite score might deprioritise weight-management or nutritional programs at a time when BMI-based screening signals rising abnormality.
Real-world parallel: imagine a university that reports an improved institutional composite score after a season of emphasis on sprint and power training. If BMI prevalence data are not reported alongside, rising overweight and underweight prevalence could go unnoticed until more serious cardiometabolic or psychosocial problems emerge.
Why composite scores can mask weight trends
Two structural features of composite surveillance explain the observed divergence.
-
Weighted composition of the total score. BMI contributes only 15% to the NSPFS total. Improvements in high-weighted components—endurance/strength (30%) and sprinting (20%)—can offset adverse shifts in BMI. For instance, better coordination, increased practice time, or coaching focusing on test-specific skills can elevate component scores without altering underlying body composition.
-
Structural overlap but non-equivalence. BMI is both a component of the total and an independent population-level distributional endpoint. That partial overlap implies some structural concordance: when BMI shifts dramatically, the composite will reflect part of that movement. But because the composite integrates several, differently weighted components with independent dynamics, the composite and BMI distribution remain non-interchangeable surveillance signals.
Measurement and administration factors further complicate interpretation. Testing conditions (equipment, examiner assignment, student motivation) and scoring calibration may change year to year. The surveillance database lacked complete round-by-round records of equipment models, calibration, examiner assignment and training, ambient conditions, and student effort. Calendar-year fixed effects account for average differences across rounds but cannot correct non-equivalent measurement. For example, heavier terminal-digit heaping in 2022 height/weight records suggested lower measurement precision that year and could partly explain the 2022 composite trough.
A practical takeaway: surveillance reports should present component-level trends and BMI categories alongside the composite. That disaggregation reveals whether observed composite changes arise from broad improvements across health-related capacities or from concentrated gains on particular tests that may not reflect holistic health improvement.
Methodological considerations, diagnostics and robustness checks
A rigorous analytic approach underpinned the findings.
-
Multilevel modelling: Random-intercept and random-slope specifications accounted for class-year clustering and allowed the within-class BMI–fitness slope to vary across units. The random-slope variance for the within-class BMI coefficient was modest (SD ≈ 0.27), with almost all class-specific slopes negative; fewer than 4% of units had a positive slope. The non-significant intercept–slope covariance argued against retaining an intercept–slope correlation but did not invalidate the random-slope specification.
-
Sensitivity analyses: Models were refitted within each survey year, confirming that between-class variation and the direction of core associations (within-class BMI, mean class BMI, and sex) were consistent across years. Cluster-level influence analyses—excluding 14 class-year units with extreme random intercepts—left core coefficients essentially unchanged.
-
Data-quality checks: The 2024 increase in BMI-defined abnormality was scrutinised. Terminal-digit heaping analysis showed that 2022—not 2024—displayed anomalous heaping in height and weight recording. Exact duplicates and constant-value classes were absent. Excluding very-low-weight records in 2024 changed prevalence negligibly. The 2024 increase was broadly distributed across classes, sexes and grades. These checks reduce the likelihood that the observed rise in abnormality is an artefact of a small number of faulty measurements.
-
Limitations inherent to the surveillance design: Stable cross-year identifiers were unavailable, so the study used repeated cross-sections rather than true longitudinal tracking. Consequently, calendar-year and grade contrasts are between-pool comparisons and cannot disentangle period, cohort, and administration effects. Class-level predictors are compositional aggregates, not directly measured psychosocial or contextual constructs. BMI is a limited anthropometric measure and cannot by itself indicate metabolic health or body composition nuances (lean vs fat mass, central adiposity).
Taken together, diagnostics and sensitivity analyses support the robustness of the main descriptive patterns, while the limitations temper causal interpretation.
Practical implications for universities and health practitioners
The study yields immediate, actionable implications for how universities conduct, analyse and report fitness surveillance.
-
Use multilevel or cluster-robust methods in analysis. With an ICC near 0.18, single-level regressions will understate standard errors and exaggerate the precision of individual-level associations. Institutions should account for class or other administrative clustering in routine analytics.
-
Report disaggregated indicators. Surveillance outputs should present composite scores together with component-level results (endurance, sprint, flexibility, etc.) and BMI-category prevalences, preferably standardised to the institution’s sex and grade distribution. Decision-makers need both composite and component-level information to target interventions effectively.
-
Interpret composite-score changes cautiously. A rising composite score does not automatically mean population-level health improvements. When composites and BMI categories diverge, immediate data-quality checks and component-level reviews should be undertaken before policy conclusions.
-
Target interventions on measured need. Class-level screening might identify groups with high mean BMI or low endurance. Interventions should be based on directly measured outcomes rather than on compositional proxies such as sex ratio, to avoid stigmatisation and to focus resources where they will reduce risk.
-
Maintain consistent testing protocols and metadata. Surveillance utility depends on measurement equivalence across rounds. Institutions should maintain records of equipment, examiner training, testing conditions and make-up testing arrangements. These metadata enable analysts to separate biological or behavioral change from administration-driven variation.
Example policy action: If a university observes rising overweight prevalence alongside stable composite scores, the institution might establish targeted nutrition counselling, expand access to strength-and-cardio training, and deploy class-level programming designed to alter peer norms around physical activity, rather than concluding fitness is improving and reducing prevention efforts.
Recommendations for researchers and practitioners
Researchers and institutional practitioners can adopt specific practices to improve inference and decision-making:
-
Collect and retain stable, ethically governed cross-year identifiers. Longitudinal linkage enables true within-student trajectory modelling and the separation of cohort, period and administration effects.
-
Add class-level psychosocial measures to routine surveillance. Simple, validated instruments measuring motivational climate, peer norms, collective efficacy, instructor engagement and facility access would allow testing whether compositional associations reflect contextual processes.
-
Expand surveillance beyond BMI. Integrate body-composition measures (e.g., bioelectrical impedance, waist circumference) where feasible to distinguish fat mass from lean mass and better characterise cardiometabolic risk.
-
Pre-register analysis plans for institutional surveillance reporting. Pre-specified analytic workflows reduce post hoc rationalisation when composite and component signals diverge.
-
Use component-weight sensitivity checks. Because the composite depends on fixed weights, sensitivity analyses that vary component weights can reveal how robust a reported trend is to alternative scoring priorities.
-
Present uncertainty. Report confidence intervals, and where clustering exists, provide cluster-robust intervals and effect sizes that reflect hierarchical structure.
These recommendations translate the study’s diagnostic findings into improved practice that preserves comparability, reduces misinterpretation and focuses interventions.
Future research directions
The study highlights several paths for further work:
-
Longitudinal multilevel models. Ethical, governed linkage of students across years would permit mixed-effects models with student-level random effects, isolating within-student change from cohort and administrative confounding.
-
Mechanism-focused class-level measurement. Direct measurement of class psychosocial and structural variables would test whether peer norms, instructor characteristics or facility constraints account for between-class variance.
-
Multi-institutional comparisons. Extending analysis to multiple universities and provinces would evaluate generalisability and allow modelling of higher levels (major, college, university) in the hierarchy.
-
Component-specific surveillance. Separate analyses of endurance, strength, speed and flexibility components could reveal whether particular capacities are improving while others lag, and whether those patterns correlate differently with BMI or class composition.
-
Intervention trials at the class level. Randomised or quasi-experimental interventions implemented at the class level would test whether altering class environments (schedules, PE pedagogy, social-motivational components) reduces between-class variation and improves individual outcomes.
Research that follows these directions will permit policy that is both evidence-based and context-sensitive.
The immediate takeaway for institutional surveillance
Class-level clustering in university fitness is substantial and consequential. Composite fitness scores provide valuable summary information, but they are not sufficient on their own. Rising prevalence of BMI-defined abnormality in the presence of stable or improving composite totals signals a disconnect between different surveillance dimensions. Institutions should adopt multilevel analytic practices, report component-level and BMI-category information alongside composite scores, retain measurement metadata, and, where possible, move toward ethically governed longitudinal tracking. These steps will produce surveillance outputs that better reflect student health needs and guide targeted, effective interventions.
FAQ
Q: Does a higher composite fitness score always mean students are healthier? A: No. The composite score aggregates multiple components with predefined weights; BMI contributes only 15% of the total. Improvements in speed, power or flexibility can raise the composite while the population BMI distribution worsens. Health is multidimensional; composite scores should be interpreted alongside BMI categories and component-specific outcomes.
Q: What does an ICC of 0.178 mean in practice? A: It means roughly 18% of the variability in composite fitness scores is attributable to which class-year unit a record belongs to. Practically, outcomes correlate within classes, and analyses must account for that clustering to avoid overstating precision.
Q: Why were students not followed longitudinally? A: The surveillance data did not include stable, ethically shareable cross-year identifiers; identifier-assignment rules differed across cohorts and years. The institutional ethics approval and data-protection agreements precluded reconstructing individual-level longitudinal links. Consequently, the analysis used repeated cross-sections rather than within-student longitudinal models.
Q: Is the rise in BMI-defined abnormality proof of deteriorating student health? A: Not by itself. BMI is an anthropometric screening tool with limitations. It does not distinguish fat from lean mass, nor does it capture central adiposity or metabolic markers. The reported rise indicates a shift in the weight-status distribution recorded under NSPFS cut-points but should be interpreted cautiously and complemented by body-composition and metabolic indicators where feasible.
Q: Could measurement or administration changes explain the year-to-year patterns? A: Measurement and administration differences can influence recorded outcomes. The analysis included calendar-year fixed effects and conducted data-quality checks (e.g., terminal-digit heaping, duplicate records). While some anomalies were present in 2022 (height/weight heaping), the steady rise in BMI-defined abnormality across years and the distribution of changes across many classes and both sexes suggest the trend is unlikely to be explained solely by measurement artefacts.
Q: What should universities change about their fitness surveillance reports? A: Report composite scores with component breakdowns and BMI-category prevalences; use multilevel or cluster-robust statistical methods; include measurement metadata with each round; and, if feasible, collect class-level psychosocial measures and ethically governed identifiers for longitudinal tracking.
Q: Would changing the weighting of the composite score affect conclusions? A: Potentially. Because the composite is a weighted sum, altering weights changes its sensitivity to component trends. Sensitivity analyses that vary weights can reveal whether observed institutional trends are robust to different priorities (e.g., placing greater emphasis on BMI or endurance).
Q: Are the findings generalisable beyond this single university? A: The analysis covers a large sample within one institution. While the patterns identified—between-class clustering and potential composite/BMI divergence—are conceptually general, the magnitude of effects and trends may differ across regions, institutions and testing practices. Multi-institutional studies are needed to establish broader generalisability.
Q: If I manage a university PE program, what is the first practical step I should take? A: Begin by ensuring your surveillance reports display component-level results and BMI-category prevalences alongside the composite. Review class-level aggregates to identify units with high mean BMI or low endurance. Simultaneously, document testing procedures, equipment and examiner assignments to improve comparability across rounds.
Q: How can researchers build on this work? A: Secure ethically governed identifiers to enable longitudinal models; incorporate direct measures of class-level psychosocial and structural variables; expand surveillance to multiple institutions; and design class-level intervention studies to test whether modifying class environments reduces between-class variation and improves outcomes.