Gender gaps in student performance,
seen in context
Explore how gaps at the average and at the top-decile threshold vary across economies. Compare indicators, inspect individual countries, and check how the picture changes with the sample.
P90: the cutoff for the top 10% within each gender, not the average within that top 10%. Average: all assessed students within each gender.
Overall (constructed): the equal-weight average of the three subject gaps. It is not an official PISA composite; opposing subject gaps can cancel.
Bottom 10% · P10: the 10th-percentile score within each gender: the boundary of its lowest-scoring 10%, not the average score of that group. Select P10 below, or compare all three attainment measures. The sign convention remains boys minus girls.
New PISA benchmark: the horizontal axis can show the published mean for all assessed students in each economy, with girls and boys pooled. Choose the three-subject mean or an individual subject. The vertical axis continues to show the gender gap at the average, P90 or P10. The three-subject mean is constructed here and is not an official PISA composite.
Choose countries one at a time. Click a name below to remove it. Highlighting does not change the correlation sample.
JPEG downloads use the displayed language and settings. Long tables are split into numbered JPEGs in one ZIP file.
Interactive scatterplots
How to read a pattern
A positive correlation means the signed boys-minus-girls gap tends to become more positive as the selected indicator rises. A negative correlation means it tends to become more negative. Neither automatically means that overall gender equality improves or deteriorates.
Pearson r describes linear association. Spearman ρ compares ranks and is useful for checking whether the pattern depends on extreme values. Values near zero indicate little association of that particular kind.
A larger coefficient in one view is not by itself evidence that the population relationship is stronger. The paired average-versus-P90 comparisons below show the uncertainty in that difference.
Try: compare mathematics at average and P90; switch GDP between dollars and logs; hold the countries fixed when moving between indicators. Each of these asks a different question.
The points are economies with equal analytical weight. These charts do not control for income distribution, region, school systems or other possible explanations. “Overall” is an analyst-defined average of three signed subject gaps; opposite gaps can cancel.
Compare correlations across indicators
The table uses signed gaps and responds to the sample and country omissions above. PISA all-student performance comes first, followed by HCI, GDP and WEF. GDP rows use the GDP scale control; the other indicators use their original scales. Select a value to open its charts. Sample sizes appear beside each coefficient.
The shared-country filters retain the original WEF 2025 + HCI + GDP overlap (79 economies before subject availability and omissions). They do not require complete pooled PISA scores: the new three-subject benchmark still requires all three subject means. Each cell shows its actual sample size. Use the same economies when comparing associations.
WEF education: inspect the influence of a few economies
Of the 81 matched economies with a WEF education score, six have a published score below 0.97: Albania (0.955), Cambodia (0.953), Kenya (0.847), Morocco (0.960), Rwanda (0.960) and Zambia (0.930). The other scores cluster near 1.000. The comparison below shows how much the coefficients change when selected observations are omitted.
These are post-hoc sensitivity checks, not recommended exclusions. The full sample remains the reference. Removing the six economies narrows the range of education scores and changes the population being compared. Dependence on these observations does not establish a data error or a causal explanation.
This table always uses signed gaps and the full matched samples. It is independent of the main chart controls, filters and manual omissions. No new p-values or confidence intervals are assigned to these selected omissions. Δr is the comparison coefficient minus the full-sample coefficient.
Average versus P90: paired comparisons and uncertainty
Fixed full matched sample for the selected indicator and GDP scale; signed gaps. The difference is |r(P90)| − |r(average)|. Its 95% interval comes from resampling the same economies jointly. Positive values favour a stronger P90 association; negative values favour the average. These exploratory difference intervals are not adjusted for multiple comparisons.
Sampling uncertainty matters. In this workbook, P90 gap estimates generally have larger standard errors than mean gaps. The intervals below resample economies and do not propagate the uncertainty of the PISA scores within each economy. Under a classical independent-error model, measurement noise can weaken correlations, but that does not establish the direction or size of bias in these comparisons. A stronger observed P90 association cannot therefore be described as necessarily understated. Mean and P90 differences alone also do not identify score variance or its causes.
| Measure | n | Average r | P90 r | Difference in |r| | 95% interval |
|---|
P10 versus average and P90: paired comparisons and uncertainty
Fixed full matched samples for the selected indicator and GDP scale; signed gaps. Each row compares P10 with the stated reference measure on the same economies. The difference is |r(P10)| − |r(reference)|: positive values favour a stronger P10 association, negative values favour the reference. Its 95% interval jointly resamples the same economies 20,000 times. These exploratory intervals are not adjusted for multiple comparisons and do not propagate within-economy PISA sampling uncertainty. They do not replace the original average-versus-P90 comparison above.
| Measure | Reference measure | n | Reference r | P10 r | P10 difference in |r| | 95% interval |
|---|
Full-sample coefficients and multiple-comparison checks
Original average/P90 results: this retained section covers the original eight tests per indicator/scale, 40 WEF tests and 128 tests overall. Its coefficients, intervals and corrections are unchanged. The additional table below reports expanded corrections that also include P10.
These reference results use all valid matches and signed gaps, regardless of current filters or omissions. The table reports Holm correction across the eight mean/P90 tests for the selected indicator and scale, and across all original indicator/scale tests. WEF also has a correction across its available comparisons. These are exploratory analyses, not preregistered confirmatory tests. The different testing families can produce different significance labels.
Read the interval and the test together. The p-values use the conventional Pearson correlation test; the intervals use percentile resampling of economies. These are different procedures with different assumptions, so they need not give the same answer. Highlighted rows flag an eight-test Holm p-value below 0.05 alongside an interval containing zero. Treat such results as sensitive to the inference method. A small adjusted p-value does not address influential observations, sampling selection or measurement error. Pearson-test assumptions.
Holm correction does not require independence between tests, provided the individual p-values are valid. The 128-test correction counts all original specifications; closely related GDP versions make this a conservative exploratory overview. No specification has been retrospectively designated as preregistered or primary. The unadjusted bootstrap intervals and the adjusted p-values do not control the same error rate. Multiple-comparison documentation.
| Measure | Level | n | Pearson r | Spearman ρ | Bootstrap 95% interval | Holm p: 8 | Holm p: WEF | Holm p: all |
|---|
Including P10: full-sample results and expanded multiple-comparison checks
These retained full-sample signed-gap results cover the original indicator snapshots: 12 tests per indicator/scale, 60 WEF 2025 tests, and 192 tests overall. They exclude the new PISA x-axis benchmarks. Their coefficients, intervals and corrections remain unchanged. The next table adds the PISA benchmarks and a separate 240-test correction. A highlighted interval includes zero despite a 12-test adjusted p-value below 0.05. These are exploratory testing families; intervals remain unadjusted.
| Measure | Level | n | Pearson r | Spearman ρ | Bootstrap 95% interval | Holm p: 12 | Holm p: WEF 60 | Holm p: all 192 |
|---|
Including PISA performance: full-sample results and expanded checks
Fixed full matched samples and signed gaps; current filters and omissions do not alter this table. Holm corrections cover 12 tests per indicator/scale, 60 WEF 2025 tests, 48 PISA-benchmark tests (four x-axis benchmarks × four gap domains × three attainment measures), and 240 tests across all indicators/scales. Earlier coefficients and bootstrap intervals are retained exactly. Only the new PISA associations and the 240-test correction are additions. All are exploratory; the 95% bootstrap intervals remain unadjusted and do not propagate survey uncertainty.
| Measure | Level | n | Pearson r | Spearman ρ | Bootstrap 95% interval | Holm p: 12 | Holm p: WEF 60 | Holm p: PISA 48 | Holm p: all 240 |
|---|
WEF education parity is not a transformed PISA score
The WEF Educational Attainment subindex uses literacy and enrolment in primary, secondary and tertiary education. These measure access and basic attainment parity; PISA measures performance on particular assessments. WEF’s 2025 scoring framework does not list PISA scores among the four education inputs. Comparing the two is therefore a comparison of different constructs, not a reconstruction of WEF’s processing of PISA. WEF 2025 framework and education discussion.
WEF generally converts its input indicators into female-to-male ratios and caps them at parity. Ratios above 1 receive the same capped value as 1; the two health indicators have different benchmarks. This deliberately focuses on shortfalls affecting women and girls rather than measuring disadvantage in both directions symmetrically. It can produce a ceiling and many tied scores. Under this scoring rule, female disadvantage is reflected in the index, while female advantage is capped at parity. The index therefore does not measure disadvantage to women and men symmetrically. WEF user guide.
What capping removes
Illustrative literacy or enrolment ratios, not real countries and not PISA scores:
A ratio of 1.20 and a ratio of 1.00 both become 1.00. The first describes a higher female value; the second equal values. The cap discards that distinction. It does not imply that the corresponding students have equal reading, mathematics or science scores. PISA score scales do not have a meaningful absolute zero, so applying female/male score ratios to PISA would itself be inappropriate.
For P10, the signed reading gaps among the 30 matched economies with a published WEF education score of 1.000 range from -69.31 to -10.11 points; girls have the higher P10 score in 30 of these economies. This is a fixed-sample illustration and does not follow the chart filters.
Rounding matters: 41 of the 148 published education scores are 1.000 at three decimals, while 35 economies share rank 1. A printed 1.000 alone does not establish exact parity. Source images: economic, education, health, political.
WEF distinguishes index inputs from complementary indicators shown in economy profiles; the latter are not incorporated into the index. The education subindex uses the literacy and enrolment indicators described above; no PISA score enters that calculation. The comparison here does not rely on any contextual profile indicator. The “2025” WEF label is the report edition, not a guarantee that every underlying observation was collected in 2025.
GDP sources: compare the same year and countries
GDP is nominal output per resident in current US dollars. The World Bank and IMF provide 2025 values; the archived UN series ends in 2024. All three 2024 versions are included so source differences can be examined on the same GDP year. No source is filled using another, and no 2024 value is labelled 2025.
Where source values diverge
Materiality here means a symmetric percentage difference above 10%: 100 × |A − B| / ((A + B) / 2). This is a descriptive threshold, not a significance test. Country-level differences may reflect revisions, estimation, population denominators or conversion methods; no specific cause is assigned here.
Country data and exact plotted values
Missing values are left missing. The CSV preserves source years, status, full-precision scores, signed gaps and any transformed plotting values.
| Country / economy | Indicator | Year | Level | Overall | Mathematics | Science | Reading |
|---|
Sources, scope and reproducibility
Reference periods and release versions
| Indicator | Period used here | Release or snapshot |
|---|---|---|
| PISA | 2025, as labelled in the supplied workbook | Supplied Tables I.B1.2c.1–3; assessment cycle and publication date are distinct. |
| PISA all-student means | 2025 assessment; both genders pooled | OECD xgs41b.xlsx, version 1, 8 September 2026. |
| WEF | 2025 report edition; underlying observation years vary by indicator and economy | 11 June 2025 release; the edition year is not a common measurement year. |
| Classic HCI | 2020 index, total for both sexes; component inputs come from different, often earlier years | World Bank 2020 HCI release; includes earlier learning assessments. |
| World Bank GDP per capita | 2024 and 2025 | Archived WDI API snapshot updated 13 July 2026. |
| IMF GDP per capita | 2024 and 2025 | April 2026 WEO; 27 of 86 matched 2025 values are flagged as staff estimates. |
| UN GDP per capita | 2024 | January 2026 National Accounts upload. |
Matching economies does not align measurement years, cohorts or definitions. These archived snapshots do not update automatically.
PISA
Unchanged user-supplied workbook 68stqn.xlsx, labelled PISA2025, Tables I.B1.2c.1–3. Results refer to that supplied workbook. Its identity is recorded by SHA-256 in the downloadable analysis metadata. The original OECD microdata and survey fieldwork have not been independently audited here.
Girls’ mean / P90: columns B / N; boys’ mean / P90: Q / AC; published boys-minus-girls gaps: AF / AR; gap standard errors: AG / AS. Each numeric published gap is checked against score subtraction. The mean covers assessed students within each gender, not the whole national population.
Coverage and selection: PISA targets students aged 15 years 3 months to 16 years 2 months who are enrolled in grade 7 or above, subject to sampling exclusions. The share of the entire age cohort represented varies between economies. Differences in enrolment, exclusions and participation by gender could affect the observed gaps and their associations with an enrolment-based index. This is a possible source of selection, not an explanation established by these charts. Country-specific Coverage Index 3 values are not included here. OECD sampling and coverage explanation.
“Overall” is the unweighted mean of three subject gaps, requiring all three. It is not an official PISA composite or a percentile of combined scores. No Overall standard error is calculated without cross-subject covariance.
P10 addition: girls’ P10 score / SE are in columns H / I; boys’ in W / X; the published boys-minus-girls P10 gap / SE in AL / AM. All three tables explicitly label these columns “10th percentile”. Published gaps are used without rounding and checked against boys minus girls within 0.000021 points. The same 87 country matches, indicator snapshots and exclusions are retained; Uzbekistan’s mathematics and reading remain missing. “Overall” at P10 averages the three signed P10 subject gaps, requires all three and is not a combined-score percentile. No Overall SE is inferred from marginal subject SEs.
All-student benchmark: OECD workbook xgs41b.xlsx, version 1, updated 8 September 2026, from PISA 2025 Results (Volume I). Science uses Table I.2.1, reading Table I.2.o1, and mathematics Table I.2.o2, column B. These published, survey-weighted means combine the assessed girls and boys; they are not an unweighted average of the two gender means. The constructed three-subject benchmark gives the three published means equal weight and requires all three. Full available precision is retained. The same 87-economy scope is retained; Uzbekistan has only science.
Related measurements: the x-axis always uses the all-student mean, even when the gap on the y-axis is P90 or P10. Both axes come from the same assessment, so their measurement errors can be related. For the mean gap, the pooled mean and the boys-minus-girls difference also share the same component scores; unequal variation in boys’ and girls’ scores across economies can contribute to their correlation. These plots describe a pattern, not an independent explanation of its causes. The bootstrap resamples economies and does not model that within-economy measurement covariance.
WEF
Global Gender Gap Report 2025, published 11 June 2025. The overall scores were checked against official Table 1.1. All 592 component scores were extracted from the four official Table 1.3 images using two agreeing OCR passes. The mean of the four components was checked against the overall score for every economy, allowing for rounding. Source precision is three decimals; Spearman calculations retain the resulting ties.
Human Capital Index
World Bank HD.HCI.OVRL, source 63, classic HCI 2020, total for both sexes. All 174 archived country values agree between official JSON and CSV formats. This uses neither an interpolated 2025 HCI nor the newer HCI+.
The HCI view asks how PISA gender gaps relate to a country’s overall human-capital level. HCI for both sexes is not a gender-parity index. A comparison with an HCI gender gap or ratio would answer a different question and is not included in this version.
HCI already contains education and harmonized learning measures, including earlier PISA results, so it is not an education-independent benchmark. Its hypothetical newborn cohort differs from the students assessed in PISA. World Bank HCI 2020 report, appendix C3.
GDP per capita
World Bank WDI, NY.GDP.PCAP.CD: 2025 and 2024, API updated 13 July 2026. IMF April 2026 WEO, NGDPDPC: 2025 and 2024; 27 of the 86 PISA-matched 2025 values are beyond the latest-actual period and flagged as IMF staff estimates. UN National Accounts Main Aggregates: 2024, January 2026 upload. All are nominal current US dollars, not PPP income.
Country matching and uncertainty
National GDP/HCI/WEF values are not substituted for B-S-J-Z (China), Dushanbe (Tajikistan), Kurdistan Region (Iraq), or Ukrainian regions (17 of 27). This leaves 87 potentially matchable economies; actual coverage varies by indicator. Chinese Taipei is matched to its separate IMF GDP entry only. Missing PISA mathematics and reading for Uzbekistan remain missing. OECD averages are excluded.
Each economy has equal analytical weight. Pearson, tied-rank Spearman and OLS describe associations. The 20,000-resample bootstrap jointly resamples economies with recorded seeds. These intervals do not propagate PISA survey-design errors, index uncertainty or GDP measurement error, and countries may be regionally dependent. Subject error bars use ±1.96 × the published gap SE; they are not regression weights.
Gender categories follow the supplied sources. The data do not describe every gender identity, within-country variation, or individual capability. This analysis neither ranks the worth of populations nor determines which policies caused any pattern.
Reproduce and review
The original analysis bundle contains source snapshots, country mappings, Python/Bash scripts and verification reports; bash run_analysis.sh reproduces that analysis. The version 1.1 update bundle contains the frozen HCI-default page, the revision script and browser checks. Its bash run_update.sh rebuilds the version 1.1 presentation and the added sensitivity table. The original embedded data and statistical metadata are unchanged. Automated numerical and browser checks reduce some implementation risks, but do not constitute independent human review or peer review.
What changed in version 1.1?
Updated 14 September 2026 following a separate validation report. This revision adds the WEF education sensitivity comparison, flags disagreement between inference methods, clarifies coverage, measurement uncertainty and source periods, and fixes the slider label markup. HCI remains the default. Country values and the original full-sample coefficients, bootstrap intervals and multiple-comparison corrections are unchanged.
As part of the follow-up, 1,036 numeric P90 score/gap/standard-error cells and 1,036 corresponding mean cells were compared with the supplied workbook; all matched exactly. This verifies transcription against that file, not independent authentication of its origin or an audit of the survey. The additional review and automated checks do not constitute peer review or remove the disclosure at the top of this page.
What changed in version 1.2?
This version adds Japanese alongside English in the same file. The language switch changes the presentation only: data, numerical precision, statistical calculations, chart geometry and machine-readable downloads are shared. Your current controls, country selection and omissions are preserved when you switch. A fresh visit starts in English; a shared link can specify either language.
The separate bilingual support bundle contains the English v1.1 input, translation dictionaries, presentation code and parity checks. Run bash run_build.sh to rebuild this single HTML file. The original analysis and v1.1 revision evidence remain available through the existing downloads. Translation and automated checks do not replace specialist review.
Bottom 10% addition: scope and verification
Added 14 September 2026. P10 is now available in the chart selector, country details, correlation matrix, GDP source comparisons, WEF education sensitivity table and plotted CSV. The original average/P90 data and statistical metadata are retained unchanged. New P10 data, paired comparisons and expanded corrections can be downloaded separately. The P10 support bundle includes the frozen input HTML and workbook, extraction and analysis scripts, a separate verifier and rebuild instructions. Automated verification checked 1,554 P10 score/gap/SE cells against a second workbook reader, all 64 new full-sample correlations and the paired bootstrap intervals. This does not constitute independent human or specialist review, and the disclosure at the top still applies.
17 September 2026 update: all-student PISA performance is now the default and appears first in the correlation matrix. Four x-axis benchmarks and 48 correlations were added. Original gender scores, HCI/GDP/WEF 2025 values and prior statistical estimates are unchanged. The support bundle contains the frozen input HTML, OECD workbook, Python extraction and independent verification scripts, interface checks and a Bash rebuild script. These automated checks do not replace human review. WEF 2026 numerical data are still pending incorporation.