Research & Data
This section holds only Nihongo to Japan's original research and publicly accessible aggregate data — things with a stated method, a sample size, a limitations section, and a citation format. Regular learning and travel articles are not published here.
⚠️ “Publicly accessible” is not the same as “openly licensed”. These datasets are not open data and carry no CC BY / CC0 licence. The figures may be cited with attribution and a link back; please contact us before reproducing our charts.
Datasets and reports
Japanese Learning Answer Dataset — August 2026
Published 2026-08-14 · version 1.1.0 | 56,613 anonymous answer events, 1,053 anonymous learner identifiers, 11,987 questions. Covers 2–13 August 2026 (12 days) — not a full-year study.
Raw accuracy broken down by skill axis, JLPT level and 39 knowledge points, with Wilson 95% confidence intervals, a minimum cell size rule, an analysis-set decomposition and a full limitations section. Downloadable as JSON and CSV, with APA, MLA and BibTeX citation formats. ⚠️ Not openly licensed: figures may be cited with attribution; please contact us to reproduce the charts.
Headline findings (quiz-context first attempts): grammar patterns 53.7% versus vocabulary 92.1%; the particle が 60.9% versus を 78.7%.
Read the full report → | JSON | CSV
How Taiwan Learns Japanese 2026
Published 2026-07-21 | Public data from the Japan Foundation, official JLPT statistics and JNTO, cross-referenced with this site's own search and learning-behaviour data.
What it answers: how Taiwanese learners actually study Japanese — the shape of lookup demand, why learners stall on nuance rather than vocabulary, and the shrinking classroom versus growing self-study picture.
How we handle data
- Frozen snapshots: every figure in a report comes from a single frozen point in time, so the numbers do not drift between paragraphs or between days.
- Method and limitations in the open: sampling, analysis-set definitions, minimum cell size and known gaps are all stated up front, not buried in a footer.
- Aggregates only: the public files contain no raw answer events and no personally identifying data of any kind.
- No overclaiming: observational data is described as observational, correlation is described as correlation, and nothing is dressed up as a representative study or a causal finding.
- Unobservable is stated as unobservable: behaviour the instrumentation never recorded is never written up as "it did not happen".
An introduction for educators and librarians is here. For the site and its editorial standards see About and the editorial policy.