PISA · gender parity · economic output · human capital

Gender gaps in student performance,
seen in context

Explore how gaps at the average and at the top-decile threshold vary across economies. Compare indicators, inspect individual countries, and check how the picture changes with the sample.

Vertical: boys’ PISA score minus girls’ score. Positive: boys higher. Negative: girls higher.
P90: the cutoff for the top 10% within each gender, not the average within that top 10%. Average: all assessed students within each gender.

Boys higher Girls higher PISA sampling cautionDashed line: descriptive linear fitPurple: selected economy

How to read a pattern

A positive correlation means the signed boys-minus-girls gap tends to become more positive as the selected indicator rises. A negative correlation means it tends to become more negative. Neither automatically means that overall gender equality improves or deteriorates.

Pearson r describes linear association. Spearman ρ compares ranks and is useful for checking whether the pattern depends on extreme values. Values near zero indicate little association of that particular kind.

A larger coefficient in one view is not by itself evidence that the population relationship is stronger. The paired average-versus-P90 comparisons below show the uncertainty in that difference.

Try: compare mathematics at average and P90; switch GDP between dollars and logs; hold the countries fixed when moving between indicators. Each of these asks a different question.

The points are economies with equal analytical weight. These charts do not control for income distribution, region, school systems or other possible explanations. “Overall” is an analyst-defined average of three signed subject gaps; opposite gaps can cancel.

Compare correlations across indicators

The table uses signed gaps and responds to the sample and country omissions above. GDP rows use the GDP scale control; WEF and HCI use their original scales. Select a value to open its charts. Sample sizes appear beside each coefficient.

Use the shared-country sample to hold membership fixed across the three indicator families. Coefficients on different samples are not directly comparable evidence of a difference in association.

Average versus P90: paired comparisons and uncertainty

Fixed full matched sample for the selected indicator and GDP scale; signed gaps. The difference is |r(P90)| − |r(average)|. Its 95% interval comes from resampling the same economies jointly. Positive values favour a stronger P90 association; negative values favour the average. These exploratory difference intervals are not adjusted for multiple comparisons.

MeasurenAverage rP90 rDifference in |r|95% interval
Full-sample coefficients and multiple-comparison checks

These reference results use all valid matches and signed gaps, regardless of current filters or omissions. The table reports Holm correction across the eight mean/P90 tests for the selected indicator and scale, and across all available indicator/scale tests in this page. WEF also has a correction across its available comparisons. These are exploratory analyses, not preregistered confirmatory tests. The different testing families can produce different significance labels.

MeasureLevelnPearson rSpearman ρBootstrap 95% intervalHolm p: 8Holm p: WEFHolm p: all
Different measurements, different questions

WEF education parity is not a transformed PISA score

The WEF Educational Attainment subindex uses literacy and enrolment in primary, secondary and tertiary education. These measure access and basic attainment parity; PISA measures performance on particular assessments. WEF’s 2025 scoring framework does not list PISA scores among the four education inputs. Comparing the two is therefore a comparison of different constructs, not a reconstruction of WEF’s processing of PISA. WEF 2025 framework and education discussion.

WEF generally converts its input indicators into female-to-male ratios and caps them at parity. Ratios above 1 receive the same capped value as 1; the two health indicators have different benchmarks. This deliberately focuses on shortfalls affecting women and girls rather than measuring disadvantage in both directions symmetrically. It can produce a ceiling and many tied scores. That is a property of the index definition; a different PISA pattern alone does not establish an error or intentional distortion. WEF user guide.

What capping removes

Illustrative literacy or enrolment ratios, not real countries and not PISA scores:

Ratio before cap1.20
Value after parity cap1.00

A ratio of 1.20 and a ratio of 1.00 both become 1.00. The first describes a higher female value; the second equal values. The cap discards that distinction. It does not imply that the corresponding students have equal reading, mathematics or science scores. PISA score scales do not have a meaningful absolute zero, so applying female/male score ratios to PISA would itself be inappropriate.

Rounding matters: 41 of the 148 published education scores are 1.000 at three decimals, while 35 economies share rank 1. A printed 1.000 alone does not establish exact parity. Source images: economic, education, health, political.

WEF distinguishes index inputs from complementary indicators shown in economy profiles; the latter are not incorporated into the index. This page does not claim to have authenticated the year or transformation of every contextual PISA indicator in the 2025 profiles. The “2025” WEF label is the report edition, not a guarantee that every underlying observation was collected in 2025.

GDP sources: compare the same year and countries

GDP is nominal output per resident in current US dollars. The World Bank and IMF provide 2025 values; the archived UN series ends in 2024. All three 2024 versions are included so source differences can be examined on the same GDP year. No source is filled using another, and no 2024 value is labelled 2025.

Where source values diverge

Materiality here means a symmetric percentage difference above 10%: 100 × |A − B| / ((A + B) / 2). This is a descriptive threshold, not a significance test. Country-level differences may reflect revisions, estimation, population denominators or conversion methods; no specific cause is assigned here.

Country data and exact plotted values

Missing values are left missing. The CSV preserves source years, status, full-precision scores, signed gaps and any transformed plotting values.

Country / economyIndicatorYearLevelOverallMathematicsScienceReading

Sources, scope and reproducibility

PISA

Unchanged user-supplied workbook 68stqn.xlsx, labelled PISA2025, Tables I.B1.2c.1–3. Results refer to that supplied workbook. Its identity is recorded by SHA-256 in the downloadable analysis metadata. The original OECD microdata and survey fieldwork have not been independently audited here.

Girls’ mean / P90: columns B / N; boys’ mean / P90: Q / AC; published boys-minus-girls gaps: AF / AR; gap standard errors: AG / AS. Each numeric published gap is checked against score subtraction. The mean covers assessed students within each gender, not the whole national population.

“Overall” is the unweighted mean of three subject gaps, requiring all three. It is not an official PISA composite or a percentile of combined scores. No Overall standard error is calculated without cross-subject covariance.

WEF

Global Gender Gap Report 2025, published 11 June 2025. The overall scores were checked against official Table 1.1. All 592 component scores were extracted from the four official Table 1.3 images using two agreeing OCR passes. The mean of the four components was checked against the overall score for every economy, allowing for rounding. Source precision is three decimals; Spearman calculations retain the resulting ties.

Human Capital Index

World Bank HD.HCI.OVRL, source 63, classic HCI 2020, total for both sexes. All 174 archived country values agree between official JSON and CSV formats. This uses neither an interpolated 2025 HCI nor the newer HCI+.

HCI already contains education and harmonized learning measures, including earlier PISA results, so it is not an education-independent benchmark. Its hypothetical newborn cohort differs from the students assessed in PISA. World Bank HCI documentation.

GDP per capita

World Bank WDI, NY.GDP.PCAP.CD: 2025 and 2024, API updated 13 July 2026. IMF April 2026 WEO, NGDPDPC: 2025 and 2024; 27 of the 86 PISA-matched 2025 values are beyond the latest-actual period and flagged as IMF staff estimates. UN National Accounts Main Aggregates: 2024, January 2026 upload. All are nominal current US dollars, not PPP income.

Country matching and uncertainty

National GDP/HCI/WEF values are not substituted for B-S-J-Z (China), Dushanbe (Tajikistan), Kurdistan Region (Iraq), or Ukrainian regions (17 of 27). This leaves 87 potentially matchable economies; actual coverage varies by indicator. Chinese Taipei is matched to its separate IMF GDP entry only. Missing PISA mathematics and reading for Uzbekistan remain missing. OECD averages are excluded.

Each economy has equal analytical weight. Pearson, tied-rank Spearman and OLS describe associations. The 20,000-resample bootstrap jointly resamples economies with recorded seeds. These intervals do not propagate PISA survey-design errors, index uncertainty or GDP measurement error, and countries may be regionally dependent. Subject error bars use ±1.96 × the published gap SE; they are not regression weights.

Gender categories follow the supplied sources. The data do not describe every gender identity, within-country variation, or individual capability. This analysis neither ranks the worth of populations nor determines which policies caused any pattern.

Reproduce and review

The companion ZIP contains source snapshots, country mappings, Python/Bash scripts and verification reports. Run bash run_analysis.sh to rebuild using archived inputs. Automated numerical and browser checks reduce some implementation risks, but do not constitute independent human review or peer review.