Strongest result in this test
Grok 4.6
It performed the best on the whole Dataset, balancing the share of mistakes it caught with the share of useful changes it made. Overall it cost around $8.85 to run for 1000 Sentences.
A practical result for Korean AI corrections
A test of AI models on 1,418 sentences written by Korean learners: API, Assistants or Local models.
Free AI Assistant Option
DeepSeek v4.1 flash placed second in the benchmark and is likely the model powering DeepSeek chat.
Best paid option
Grok 4.6 clearly outperforms every other option - it is paid only however.
Best offline option
Meta recently released Muse Glimmer which can comfortably compete with the average paid option. At Q4,K,M it sits at 19GB in size.
The best choice depends on your goal. This benchmark can guide you from privacy and independence with a self hosted Muse Glimmer, to quality-price-complexity tradeoff with the cheap Gemini 3.5 Flash Lite.
Strongest result in this test
It performed the best on the whole Dataset, balancing the share of mistakes it caught with the share of useful changes it made. Overall it cost around $8.85 to run for 1000 Sentences.
Cheapest hosted option
The cost of running the benchmark once on a model via the API has a huge range. Google manages to take the bottom and the lead with its very cheap Gemini 3.5 Flash Lite and its very expensive Gemini 3.1 Pro. Cost can scale double with performance: Tokens get more expensive, but the model also produces more reasoning tokens.
The Price-Quality Frontier
Gemini 3.5 Flash Lite, MiMo V2.5, GLM 5.3 Flash, GPT-5.6 Terra, Claude Sonnet 5, DeepSeek V4.1 Flash, Grok 4.6
Each one gives you more measured quality than every cheaper option tested.
Each dot is a hosted model. Moving right costs more; moving up means more reliable corrections. Offline models have no API bill and so no place on this axis; the best of them is drawn as a line across the chart.
The dashed line is the best local model tested. 6 hosted models you pay for score below it.
The highlighted line connects models that offer more quality than every cheaper model tested. It is a guide to trade-offs, not a universal ranking.
Compare every completed and unfinished test. Select a column heading to sort the results.
Showing 17 runs
| Service | ||||||||
|---|---|---|---|---|---|---|---|---|
| Grok 4.6 | OpenRouter | 66.9% | 65.2% | 74.5% | 52.8% | $8.85 | 1,417 | 1,423 |
| DeepSeek V4.1 Flash | OpenRouter | 64.4% | 62.7% | 72.2% | 51.6% | $2.82 | 1,412 | 2,028 |
| Kimi K3 | OpenRouter | 63.8% | 62.1% | 71.7% | 52.2% | $9.24 | 1,417 | 622 |
| Claude Sonnet 5 | OpenRouter | 63% | 61.1% | 72.5% | 49.4% | $1.11 | 1,418 | 99 |
| Claude Opus 5 | OpenRouter | 60.7% | 58.1% | 74.5% | 43.1% | $5.74 | 1,418 | 218 |
| Qwen3.8 Max 0902 | OpenRouter | 60.4% | 58.2% | 71.1% | 48% | $11.73 | 1,391 | 985 |
| GPT-5.6 Terra | OpenRouter | 60.3% | 58% | 71.9% | 44.9% | $1.07 | 1,418 | 50 |
| GLM 5.3 Flash | OpenRouter | 60% | 58.2% | 68.6% | 46% | $0.86 | 1,412 | 995 |
| Gemini 3.1 Pro Preview | OpenRouter | 59.5% | 57.2% | 70.8% | 47.7% | $12.94 | 1,417 | 1,070 |
| GLM 5.3 | OpenRouter | 59.4% | 57.2% | 70.5% | 43.4% | $6.11 | 1,412 | 1,289 |
| Gemini 3.7 Flash | OpenRouter | 58.8% | 56.1% | 73% | 41% | $2.16 | 1,418 | 570 |
| MiMo V2.5 | OpenRouter | 58.1% | 56.7% | 64.6% | 45.2% | $0.11 | 1,417 | 297 |
| Gemini 3.5 Flash Lite | OpenRouter | 57.6% | 55.3% | 68.8% | 41% | $0.04 | 1,418 | 14 |
| MiniMax M3 | OpenRouter | 56.1% | 54.6% | 63.5% | 42.3% | $0.75 | 1,408 | 523 |
| GPT-5.6 Sol | OpenRouter | 56.1% | 53.5% | 69.9% | 40.8% | $1.09 | 1,418 | 101 |
| Qwen3.8 Flash | OpenRouter | 56% | 54.3% | 64% | 44.8% | $0.97 | 1,386 | 2,092 |
| GPT-5.6 Luna | OpenRouter | 54.5% | 51.8% | 69.2% | 37.9% | $0.16 | 1,418 | 130 |
Showing 18 runs
These have no hosted-service bill, so the price column is replaced by the weights that were actually loaded. A smaller file is quicker and fits more machines; it is also a compressed version of the model, which costs accuracy.
| Muse Glimmer | BF16 | 58.5% | 56.6% | 67.6% | 43.6% | 1,411 | 674 |
|---|---|---|---|---|---|---|---|
| Muse Glimmer | Q4_K_M | 57.9% | 55.9% | 67.4% | 41.7% | 1,418 | 693 |
| Mi:dm 2.0 Base Instruct | BF16 | 56.4% | 57.8% | 51.5% | 44.2% | 1,418 | 12 |
| Kanana 2 30B A3B Instruct | BF16 | 53.6% | 52.2% | 60.1% | 40.1% | 1,418 | 12 |
| Gemma 4 26B A4B | Q4_K_M | 52.9% | 50.4% | 66.1% | 37.4% | 1,410 | 1,011 |
| Gemma 4 26B A4B | BF16 | 52% | 49.4% | 65.8% | 32.2% | 1,418 | 15 |
| Gemma 4 12B | BF16 | 49.1% | 46.4% | 64.1% | 26.1% | 1,418 | 16 |
| Command A+ 05-2026 | FP8 | 46.5% | 44.4% | 57.3% | 30.7% | 1,415 | 359 |
| Qwen3.6 35B A3B | FP8 | 45.4% | 43.4% | 55.7% | 31.2% | 1,412 | 1,682 |
| Nemotron 3.5 Lightning 30B A3B | NVFP4 | 43% | 41.8% | 48.7% | 30.4% | 1,410 | 1,676 |
| Gemma 4 E4B | Q4_K_M | 40.7% | 38.2% | 54.8% | 16.1% | 1,418 | 81 |
| GPT-OSS 20B | MXFP4 | 39.8% | 37.8% | 50.7% | 20.5% | 1,418 | 92 |
| Kanana 2 30B A3B Thinking | Q4_K_M | 36.6% | 34.8% | 46.6% | 19.5% | 1,418 | 924 |
| Ministral 3 3B | Q8_0 | 28% | 26.3% | 37.9% | 3.7% | 1,418 | 16 |
| Ministral 3 14B Reasoning | Q4_K_M | 26.5% | 25.4% | 31.5% | 9.7% | 1,417 | 15 |
| Nemotron 3 Nano 4B | Q8_0 | 23.6% | 22.6% | 28.4% | 8.9% | 1,416 | 902 |
| DeepSeek R1 0528 Qwen3 8B | Q4_K_M | 23.2% | 21.9% | 30.3% | 4.3% | 1,404 | 382 |
| Kanana 2 3B Instruct | Q6_K | 12.5% | 12.1% | 14.1% | 0.8% | 1,418 | 13 |
Unfinished tests stay visible for transparency but are left out of the recommendations. Models for your own computer were run through LM Studio or on rented Vast.ai GPUs; their hardware, electricity, and time were not measured.
Output lengthCompletion tokens / sentence is the average generated text per answered sentence. Models used their normal settings, so some spent much more time thinking about the same one-sentence task. That affects both price and speed.
Two models can have similar scores while giving very different advice. Pick a maker and model to see whether it usually gets the sentence right, changes too much, or misses the mistake.
Change nothing at all: 40.5% fully correct (574 of 1,418 sentences)
No extra change and no missed mistake
Added a change that was not needed
Left a needed correction undone
Changed too much and still missed something
No answer was returned after retries
The details for readers who want to check the work: the sentences used, the question each model received, how answers were judged, and what this test cannot tell us.
I used KoLLA_multi-refs.m2 from the KoLLA v2 dataset:
100 learner essays, split into
1,418 sentences. They contain
2,828 human corrections
and 3,649 individual edits.
All but eight sentences have two human references.
f08943bedab1b149 9a6f2e3fea1b39bbb7343445db1167f7Every model saw the same one-line instruction:
Correct the Korean sentence. Reply with the corrected sentence only.
Hosted models ran through OpenRouter. Offline models ran through LM Studio. Temporary failures were retried up to three times. Models kept their normal response settings; I did not force the same amount of internal reasoning on every model.
The scorer compares each answer with both human corrections for that sentence. If either comparison found a valid correction, it keeps the better one. That is the usual approach for sentences with more than one valid answer, though it gives models the benefit of the doubt.
The scorer used here is M2 MaxMatch (Dahlmeier & Ng, 2012). It is checked on every run against the reference Python m2scorer.
Hosted costs are the recorded service bill. They reflect exact prices paid on execution, not the monthly price of a consumer subscription.
Offline models have no hosted-service bill, but I did not measure hardware, electricity, or time, so they do not get any cost attribution.
A test is complete when every one of the 1,418 sentences has either an answer or a recorded failure.
This list does not contain every model available, especially not for local runs. The selection for local models was arbitrary, based on what was available to me at the time of testing. The selectino of hosted models was based on the current available services and their popularity.
Before comparing models, it helps to know what counts as a good correction. Here is the short version, in plain language.
Two people corrected every sentence independently. When their answers differed, the corpus accepts either valid reading—because Korean often has more than one natural way to repair a sentence.
비행기 음식이 안 막였습니다.
비행기 음식을 안 먹었습니다.
비행기 음식이 안 맞았습니다.
574 of the 1,418 sentences were already correct, so repeating them unchanged is the right answer. That means a model that changes nothing gets 40.5% of the sentences fully right—but it fixes none of the mistakes.
Each model was run once, so nearby results are effectively tied. Use the table to compare broad differences, not to declare a winner by a point or two.
Feel free to reproduce this benchmark on your own. The code and dataset are open source, and the raw runs results are available for download for reevaluation.
The raw reports contain each sentence, each answer, and the detailed scores.
A benchmark can point you in the right direction. Try Elephant’s sentence analysis on Korean or read my personal telling of how and why I ran this benchmark.