Japanese Learning Answer Dataset — August 2026

In one line: aggregate statistics from 56,613 anonymous answer events on Nihongo to Japan, contributed by 1,053 anonymous learner identifiers across 11,987 questions.

⚠️ This is an August 2026 snapshot, not a study of the whole of 2026. The data covers 2–13 August 2026, twelve days in total (the aggregation pipeline went live on 2 August; the first day is partial). No longitudinal trend can be read from it. Version 1.1.0, snapshot frozen 2026-08-14.

On "learner identifiers": these are anonymous identifiers generated at random in the browser and cannot be merged across devices. One person using a phone and a laptop produces two identifiers, so 1,053 is not a verified count of distinct people — it is an upper bound. This page says "learner identifiers" throughout rather than "people".

What this is not. This is not a representative survey of Japanese learners in Taiwan or anywhere else. The sample consists of people who chose to practise on this site — a self-selected, observational sample. Every figure describes what happened in this set of answers, and cannot be extrapolated to any population. See Limitations.

1. Five main findings

1. Grammar patterns are the only axis below 60%, 38 points behind vocabulary

Among quiz-context first attempts (see Methodology for the set definition), vocabulary items were answered correctly 92.1% of the time (n=22,869, 627 learner identifiers, 95% CI 91.8%–92.5%), while grammar-pattern items reached only 53.7% (n=8,678, 532 identifiers, 95% CI 52.6%–54.7%).

The confidence intervals are far apart, so this is not sampling noise. The data shows a large gap between the two axes; it says nothing about why.

2. Grammar-pattern "accuracy" runs from 53.7% to 74.1% depending on which answers are included

Under different analysis-set definitions, observed accuracy on the grammar-pattern axis ranges from 53.7% (quiz-context first attempts, n=8,678) to 74.1% (all answer events, n=24,293). Because the events included and the learner-identifier composition both differ, the whole difference cannot be read as the product of a single counting rule.

Step by step:

StepnRaw accuracyChange vs previous
All answer events24,29374.1%
→ keep only first attempts22,63873.2%-0.87 pp
→ keep only quiz contexts (drop lesson/review)8,67853.7%-19.54 pp
(comparison) lesson-context first attempts13,93785.4%

Removing repeat attempts moves the figure by only 0.9 percentage points. Almost all the remainder comes from the difference between lesson-context and quiz-context answers.

Paired comparison on the same learner identifiers: 75 identifiers have at least 5 grammar-pattern first attempts in each context. They reach 86.6% in the lesson context (n=5,858) and 55.5% in the quiz context (n=2,055), a paired mean difference of +27.5 pp (95% CI 22.6–32.4; median +26.9). So the gap is not simply a matter of who is answering.

⚠️ A same-question comparison is not possible, so no conclusion can be drawn. Of 5,772 distinct grammar-pattern questions, only 6 appear in both contexts, and only 1 has at least 3 answers on each side. The two contexts draw on almost disjoint question sets, so the context effect cannot be separated from differences in question difficulty. The lesson-context questions may simply be easier.

What can honestly be said is this: when citing an accuracy figure from a learning platform, always ask which analysis set it came from. All four sets are in the downloadable data.

3. N3 and N2 have the lowest raw accuracy — but this data cannot separate them

Quiz-context first-attempt raw accuracy by level: N5 90.2%, N4 73.4%, N3 67.8%, N2 69.4%, N1 76.0%.

The intervals for N3 (95% CI 66.6%–69.0%) and N2 (95% CI 67.6%–71.1%) overlap, so it is not correct to say "N3 is harder than N2". What can be said is that in this sample, N2 and N3 both sit below N4, N1 and N5 on raw accuracy.

⚠️ More importantly, each level draws on a completely different question set. This ordering is not a difficulty comparison and is not evidence about which JLPT level is objectively harder.

4. The spread within particles is wider than the spread between levels

Quiz-context first-attempt accuracy on individual particles: が 60.9% (n=358, CI 55.8%–65.8%), に 62.4% (n=370), で 63.3% (n=330), versus を 78.7% (n=385, CI 74.3%–82.5%) and の 83.3% (n=120).

The intervals for が and を do not overlap; the gap is 18 percentage points.

5. One-vote-per-identifier and one-vote-per-answer differ by 7.04 points

Computed over exactly the same 372 learner identifiers, the same analysis set and the same 34,392 eligible events:

WeightingAccuracy
Event-weighted (one vote per answer)82.5%
Identifier-weighted (one vote per identifier)75.5% (median 76.2%)
WEIGHTING_EFFECT_PP7.04 pp

Cohort: Learner identifiers with at least 20 answers in the primary analysis set. Both figures share the same numerator population and denominator, so the gap reflects the weighting choice alone and is not contaminated by differences in sample composition.

The cause is repeated measures: the single most active identifier accounts for 3.7% of all answers and the top ten for 19.4%. Any study using platform data hits this, so we publish both numbers.

2. Dataset overview

ItemValue
Valid answer events56,613
Anonymous learner identifiers (upper bound on people)1,053
Distinct questions11,987
Knowledge points39
Temporal coverage2 to 13 August 2026 (12 days); the first day is partial
Snapshot cutoff2026-08-14T00:00:00Z
First-attempt events52,934
Quiz-context first attempts (primary set)37,245
Learner identifiers in the primary set688
Lesson-context first attempts (comparison set)15,741

Answers by source: challenge 23,369 (41.3%), self_study 16,782 (29.6%), placement_test 11,797 (20.8%), special_training 4,157 (7.3%), review 508 (0.9%).

Distribution of the database evidence_class column (a raw table field, not the name of an analysis set in this report): A 37,245 (65.8%), B 17,048 (30.1%), C 2,320 (4.1%). The primary analysis set is exactly evidence_class = A.

Bar chart of sample distribution by evidence grade and answer source
Figure 4: Sample distribution across all 56,613 valid answer events. Source: this dataset.

3. Accuracy by learning axis

Bar chart of quiz-context first-attempt raw accuracy by skill axis with Wilson 95% confidence intervals
Figure 1: Raw accuracy by skill axis (quiz-context first attempts, n=37,245, August 2026). Whiskers show Wilson 95% confidence intervals.
Skill axisnIdentifiersRaw accuracy95% CIFlag
Grammar patterns8,67853253.7%52.6%–54.7%OK
Particles1,98613769.1%67.1%–71.1%OK
Word forms / conjugation2,32238683.2%81.7%–84.7%OK
Idiomatic expressions1,34126186.4%84.5%–88.2%OK
Vocabulary22,86962792.1%91.8%–92.5%OK

The following axes fall below the minimum cell size (n<30 or fewer than 5 learner identifiers), so no accuracy figure is published — only the event count: Mixed / all-round (n=27, 8 identifiers), Reading comprehension (n=22, 4 identifiers).

4. Accuracy by JLPT level

⚠️ This is the section most easily misread. Each level draws on a completely different question set, and no IRT difficulty equating was performed. The figures below are the proportion of these particular questions answered correctly — not a measure of which level is harder.
Bar chart of quiz-context first-attempt raw accuracy by JLPT level with Wilson 95% confidence intervals
Figure 2: Raw accuracy by JLPT level (quiz-context first attempts, n=37,245, August 2026). Levels draw on different question sets; this is not a difficulty comparison.
JLPT levelnIdentifiersRaw accuracy95% CI
N519,39538090.2%89.8%–90.6%
N46,09931373.4%72.3%–74.5%
N35,98543267.8%66.6%–69.0%
N22,72621269.4%67.6%–71.1%
N13,04013976.0%74.4%–77.5%

The same levels across all three analysis sets:

LevelQuiz-context first (primary)Lesson-context firstFirst attempts (all)All events
N590.2% (n=19,395)83.7% (n=11,154)87.8% (n=30,495)88.2% (n=33,286)
N473.4% (n=6,099)83.1% (n=2,694)76.5% (n=8,758)76.6% (n=9,262)
N367.8% (n=5,985)82.1% (n=1,507)70.8% (n=7,522)71.1% (n=7,727)
N269.4% (n=2,726)79.2% (n=260)70.2% (n=2,985)70.5% (n=3,072)
N176.0% (n=3,040)86.5% (n=126)76.4% (n=3,174)76.5% (n=3,266)

5. Lowest-accuracy knowledge points

Only knowledge points with n≥100 and at least 5 learner identifiers are listed (23 qualify). n, level and category are shown alongside the percentage — small cells look extreme for reasons that are usually just noise.

Bar chart of the ten lowest-accuracy knowledge points with Wilson 95% confidence intervals
Figure 3: The ten lowest-accuracy knowledge points (quiz-context first attempts, n≥100 only, August 2026).
#Knowledge pointCategoryLevelnIdentifiersRaw accuracy95% CI
1Grammar patterns (N4)Grammar patternsN41,74623448.2%45.9%–50.6%
2Grammar patterns (N3)Grammar patternsN32,53639049.2%47.3%–51.2%
3Grammar patterns (N2)Grammar patternsN21,18319151.3%48.5%–54.1%
4Grammar patterns (N1)Grammar patternsN11,29911857.0%54.3%–59.7%
5Particle がParticles3588160.9%55.8%–65.8%
6Particle にParticles3707562.4%57.4%–67.2%
7Particle でParticles3308563.3%58.0%–68.3%
8Grammar patterns (N5)Grammar patternsN51,96320964.4%62.3%–66.5%
9Word forms (N3)Word formsN391329473.2%70.2%–75.9%
10Idioms (N2)IdiomsN21676276.6%69.7%–82.4%
11Particle をParticles3856178.7%74.3%–82.5%
12Word forms (N2)Word formsN227810780.2%75.1%–84.5%

6. Most consistent strengths

#Knowledge pointCategoryLevelnIdentifiersRaw accuracy95% CI
1Word forms (N1)Word formsN12967995.6%92.6%–97.4%
2Vocabulary (N5)VocabularyN516,02931194.4%94.0%–94.8%
3Vocabulary (N1)VocabularyN17039489.2%86.7%–91.3%
4Vocabulary (N4)VocabularyN42,82824389.0%87.8%–90.1%
5Idioms (N1)IdiomsN174211188.8%86.3%–90.9%
6Word forms (N4)Word formsN446114987.0%83.6%–89.8%
7Idioms (N4)IdiomsN41409186.4%79.8%–91.1%
8Idioms (N3)IdiomsN329014285.9%81.4%–89.4%

7. Observed patterns worth noting

Observation A: N1 raw accuracy sits above N2 and N3

N1 76.0% (n=3,040, 139 identifiers) is above N2 69.4% and N3 67.8%.

This is an observed pattern, not evidence that N1 is easier. At least three explanations cannot be ruled out with this data: only 139 identifiers answered N1 items versus 432 for N3, so the groups differ in composition; the question sets differ; and no ability estimation was performed. Answering "which level is genuinely harder" needs IRT equating and an anchor-item design, and this dataset has neither.

Observation B: the grammar-pattern axis is the lowest at every level

Axis × levelnIdentifiersRaw accuracy95% CI
Grammar patterns × N41,73423348.0%45.6%–50.3%
Grammar patterns × N32,53138949.2%47.2%–51.1%
Grammar patterns × N21,17919151.1%48.3%–54.0%
Grammar patterns × N11,29511857.0%54.3%–59.7%
Particles × N47669258.5%55.0%–61.9%
Grammar patterns × N51,93920864.0%61.8%–66.1%
Word forms / conjugation × N394830274.2%71.3%–76.8%
Particles × N51,22011175.8%73.3%–78.1%
Idiomatic expressions × N21676276.6%69.7%–82.4%
Word forms / conjugation × N227810780.2%75.1%–84.5%

Crossing axis with level, almost every one of the lowest cells is a grammar-pattern cell. That is more informative than the level breakdown on its own.

8. What teachers and materials designers can use this for

This section deliberately separates data from interpretation. The data is citable; the interpretation is our judgement, and you are free to disagree with it.

Data (citable)Interpretation (our judgement, not data)
Quiz-context first attempts: grammar-pattern axis 53.7% vs vocabulary 92.1% (n=8,678 and 22,869)On this platform there is a clear gap between recognising words and assembling sentences. If practice time has to be allocated, grammar patterns may have the higher marginal return.
が 60.9%, に 62.4%, で 63.3% sit well below を 78.7% and の 83.3%を and の have comparatively narrow functions, while が, に and で are heavily polysemous. Teaching the distinct functions of a single particle separately may work better than introducing particles one at a time.
N2 and N3 raw accuracy (69.4% / 67.8%) sit below N4, N1 and N5⚠️ Do not use this to rank difficulty. A more plausible reading is that the learners answering at N2/N3 differ in composition from those at other levels, rather than anything about the items.
Grammar patterns: 85.4% in the lesson context (n=13,937) vs 53.7% in the quiz context (n=8,678); paired difference +27.5 pp across 75 identifiersIn-lesson and quiz answering produce very different results — worth keeping in mind when platform practice scores are used to judge learning outcomes. ⚠️ But the two contexts use almost disjoint question sets, so we cannot say whether this is the context or simply easier questions.

9. Methodology

ItemDetail
PopulationAnswer records produced by learners using Nihongo to Japan.
SamplingSelf-selected, observational sample. Not random sampling; not population-representative.
Unit of analysisanswer event (one learner identifier answering one question once).
User countlearner identifiers = distinct anonymous identifiers. They cannot be merged across devices, so this is an upper bound on the number of people, not a verified headcount.
Metricraw accuracy = correct ÷ valid answer events. Not an ability estimate.
Confidence intervalsWilson 95% CI. Reflects sampling error only; does not correct for selection bias.

Three analysis sets

SetDefinitionn
Quiz-context first attempts
(primary analysis set)
Answers from challenge / placement_test / special_training where this is the learner identifier’s first attempt at that question, not a repeat within 24 hours, with complete question metadata.37,245
Lesson-context first attempts
(comparison set)
First attempts from self_study / review — answers given while working through the lesson material.15,741
First attempts (all contexts)Each learner identifier's first answer to each question, regardless of source.52,934
All answer eventsEvery valid answer event, including repeats.56,613
⚠️ The primary set makes no claim about hints or explanations. The fields used_hint / viewed_explanation / used_answer_key are false throughout the database, but only because no producer ever writes them — none of the answer-capture call sites passes these parameters, and the product currently has no pre-answer hint or answer-reveal feature. The constant false is a property of the instrumentation, not evidence about learner behaviour. These fields are therefore not used as analysis conditions, and no claim is made about hint use.

Independent verification of the primary set: recomputing the earliest event per (identifier, question) on the server, 99.59% of primary-set events (37,094 of 37,245) pass the first-attempt test. The 151 that do not are most likely cross-device — the same person answering on two devices, each recording a first attempt. Every event in the set has attempt_number = 1.

What is excluded

Minimum cell size

Cells with n<30 or fewer than 5 learner identifiers are SUPPRESSED: the event count is published but the accuracy is not. Cells with n≥30 and ≥5 identifiers but n<100 are flagged LOW_SAMPLE and should always be cited with n. Cells with n≥100 and ≥5 identifiers are flagged OK.

In this snapshot 23 knowledge points are OK, 12 are LOW_SAMPLE and 3 are SUPPRESSED.

Anonymisation

10. Limitations (please cite these alongside the figures)

If you read only one paragraph, read this one. This data describes how people who came to Nihongo to Japan to practise performed on Nihongo to Japan questions. It is not a picture of Japanese learners in Taiwan, it is not a measure of JLPT difficulty, and it cannot be used to compare the objective difficulty of different levels.

11. How to cite

Citation in papers, research reports, teaching materials and journalism is welcome. Please credit the source with a link back to this page.

APA (7th)

Nihongo to Japan. (2026). Japanese learning answer dataset — August 2026 (Version 1.1.0) [Data set]. https://www.nihongotojapan.com/en/research/japanese-learning-data-2026

MLA (9th)

Nihongo to Japan. Japanese Learning Answer Dataset — August 2026. Version 1.1.0, 2026, https://www.nihongotojapan.com/en/research/japanese-learning-data-2026.

BibTeX

@misc{ntj_jlad_2026,
  title = {Japanese Learning Answer Dataset --- August 2026},
  author = {{Nihongo to Japan}},
  year = {2026},
  version = {1.1.0},
  note = {Snapshot 2026-08-14; aggregate statistics from 56,613 anonymous answer events, 2--13 August 2026},
  url = {https://www.nihongotojapan.com/en/research/japanese-learning-data-2026}
}

Licence status

This is not an open dataset and is not under any standard open licence (CC BY, CC0 or similar). Nihongo to Japan is free to use, but "free to use" is not the same as "openly licensed".

Explicitly permitted:

Charts: the charts on this page are available to view. To reproduce our charts, please contact us first. The aggregate figures themselves may be cited under the terms above.

Explicitly not permitted: presenting these figures as representative of all Japanese learners in Taiwan or anywhere else; republishing the numbers with the sampling limitations removed; redistributing the dataset wholesale as your own data product.

For anything else, please get in touch via the about page.

12. Download the data

Two formats with identical content, both aggregate statistics only (no raw answer events):

Version 1.1.0 · snapshot 2026-08-14 · methodology version 1.1. Future updates will increment the version and preserve the earlier figures rather than overwriting them in place.

⚠️ For citation, use the versioned frozen dataset rather than the live aggregate endpoint.

A live aggregate endpoint also exists at /api/aggregates, returning current numbers on each call. It uses different definitions: an all-answer-events basis rather than quiz-context first attempts, and its totalUsers counts only the primary-set learner identifiers (688) rather than all 1,053. Its numbers change daily and are not suitable for citation.

For journalists

Three figures you can quote directly:

  1. Grammar-pattern accuracy 53.7% (n=8,678) versus vocabulary 92.1% (n=22,869) — a gap of 38 points.
  2. The particle が is answered correctly 60.9% of the time (n=358); を reaches 78.7% (n=385).
  3. Grammar-pattern accuracy runs from 53.7% to 74.1% depending on which answering contexts are included (lesson context 85.4%, quiz context 53.7%). ⚠️ The two contexts use almost disjoint question sets so the cause cannot be pinned down — but that is itself a story about how learning-platform data gets quoted.

Citation formats · Download the data · Contact

For researchers and educators

Structure: one row per aggregate cell, indexed by dimension (axis / jlpt / knowledge_point / axis_x_jlpt) × analysis_set (quiz_first_attempt / lesson_first_attempt / first_attempt / all_events). Each cell reports n, learners, accuracy, Wilson ci95_low/ci95_high and a flag. The JSON additionally carries the analysis-set decomposition, the paired comparison and the weighting effect.

Sampling: self-selected and observational; the unit of analysis is the answer event; a learner identifier is an anonymous UUID, so the identifier count is an upper bound on the number of people (the same person on two devices produces two identifiers, and identifiers cannot be linked).

Known methodological gaps — we think these are worth stating rather than hiding:

Download JSON · Download CSV · Citation formats · Full methodology

← Back to Research & Data | How Taiwan Learns Japanese 2026 | For educators & libraries