The models that write best in English

The AI Writing Benchmark measures how well leading AI models write in English. Responses to a shared corpus of English writing prompts are compared blind, without model names, and the verdicts are combined into a ranking for English.

Leading across all categories

The 10 best models

The ten best models for writing in English
#ModelScore
1Claude Opus 5 πŸ’‘ adaptive high93.9
2Kimi K3 πŸ’‘ enabled93.5
3Claude Fable 5 πŸ’‘ adaptive high93.1
4Claude Fable 5.1 πŸ’‘ adaptive high92.4
4GPT-6 Astra πŸ’‘ high92.4
6GLM 5.3 πŸ’‘ max90.7
7Claude Opus 4.8 πŸ’‘ adaptive high89.4
8Claude Sonnet 5 πŸ’‘ adaptive high86.4
9GPT-5.6 Sol πŸ’‘ high85.9
10GLM 5.3 Flash πŸ’‘ high83.8

This score is not a percentage grade: 50 means an estimated one-in-two chance of beating the average model in English.

Last updated:

See the full leaderboard

The 10 best local variants

Estimated VRAM is the approximate graphics memory needed to load the entire model.

The ten best exact variants run locally
#ModelScoreEstimated VRAM
1Muse Glimmer 30B K-Quant Dynamic πŸ’‘ high 🏠76.7β‰ˆ 32 GB
2Qwen3.8 27B Q4_K_M πŸ’‘ xhigh** 🏠66.5β‰ˆ 24 GB
3Muse Glimmer 30B K-Quant 17GB πŸ’‘ high 🏠64.2β‰ˆ 24 GB
4Qwen3.6 27B Q8_0 πŸ’‘ on* 🏠62.4β‰ˆ 32 GB
5Qwen3.8 27B Q8_0 πŸ’‘ xhigh** 🏠61.1β‰ˆ 48 GB
6Gemma 4 26B A4B Q6_K πŸ’‘ on* 🏠57.6β‰ˆ 32 GB
7Gemma 4 31B Q6_K πŸ’‘ on* 🏠57.4β‰ˆ 32 GB
8Qwen3.8 27B Q6_K πŸ’‘ xhigh** 🏠56.4β‰ˆ 32 GB
9Gemma 4 31B Q8_0 πŸ’‘ on* 🏠56.1β‰ˆ 48 GB
10Qwen3.6 27B Q4_K_M πŸ’‘ on* 🏠55.7β‰ˆ 24 GB

Some software can keep part of the model in system memory. This reduces the VRAM required, but usually makes generation slower.

90models compared
54586published AI verdicts

The ranking evaluates writing quality alone. Prompts are written directly in the language being tested, and model names remain hidden until the verdict so reputation cannot influence the evaluation. This site is a personal project. No model provider pays me, and I have no reason to favour one model over another.

Explore the results

Does AI Write Better in English With or Without Reasoning?

In the current direct sample, reasoning-enabled responses earned 66.2% of the available points in English. The uncertainty interval still includes parity, the result differs sharply from one model to the next, and no human votes have compared the two settings directly yet.