Research & Data
This section holds only Nihongo to Japan's original research and publicly accessible aggregate data — things with a stated method, a sample size, a limitations section, and a citation format. Regular learning and travel articles are not published here.
The data is released under CC BY 4.0. The aggregate data published here is free to use, including commercially, provided you give credit. ⚠️ The licence covers the data only — the charts we produce and the text of the reports remain rights-reserved; please contact us before reproducing a chart.
Datasets and reports
Japanese Learning Answer Dataset — August 2026
Published 2026-08-14 · version 1.1.0 | 56,613 anonymous answer events, 1,053 anonymous learner identifiers, 11,987 questions. Covers 2–13 August 2026 (12 days) — not a full-year study.
Raw accuracy broken down by skill axis, JLPT level and 39 knowledge points, with Wilson 95% confidence intervals, a minimum cell size rule, an analysis-set decomposition and a full limitations section. Downloadable as JSON and CSV, with APA, MLA and BibTeX citation formats. Data released under CC BY 4.0 (free to use with attribution, including commercially); ⚠️ the licence covers the data only — please contact us before reproducing our charts.
Headline findings (quiz-context first attempts): grammar patterns 53.7% versus vocabulary 92.1%; the particle が 60.9% versus を 78.7%.
For journalists and researchers: ten quotable findings (each self-contained, with n and confidence intervals) | limitations | citation formats | contact
Read the full report → | JSON | CSV
How Taiwan Learns Japanese 2026
Published 2026-07-21 | Public data from the Japan Foundation, official JLPT statistics and JNTO, cross-referenced with this site's own search and learning-behaviour data.
What it answers: how Taiwanese learners actually study Japanese — the shape of lookup demand, why learners stall on nuance rather than vocabulary, and the shrinking classroom versus growing self-study picture.
How we handle data
- Frozen snapshots: every figure in a report comes from a single frozen point in time, so the numbers do not drift between paragraphs or between days.
- Method and limitations in the open: sampling, analysis-set definitions, minimum cell size and known gaps are all stated up front, not buried in a footer.
- Aggregates only: the public files contain no raw answer events and no personally identifying data of any kind.
- No overclaiming: observational data is described as observational, correlation is described as correlation, and nothing is dressed up as a representative study or a causal finding.
- Unobservable is stated as unobservable: behaviour the instrumentation never recorded is never written up as "it did not happen".
An introduction for educators and librarians is here. For the site and its editorial standards see About and the editorial policy.