AI college entrance examination essay writing competition finals: DeepSeek V4 narrative essay won the championship with 46 points, Hunyuan's "own people" scored full marks
GPT-5.5/Fable-5/DeepSeek V4/Hunyuan 3 Preview Co-written in the 2026 Beijing College Entrance Examination essay, DeepSeek V4's narrative essay won the championship with an average score of 46 points. Hunyuan's "one's own people judge one's own people" scored a perfect score and attracted heated discussion.
AI College Entrance Examination Essay Writing Finals: DeepSeek V4 won the championship in narrative writing, and Hunyuan’s perfect score caused controversy
During the 2026 college entrance examination, Anthropic's Mythos-level large-scale models were released simultaneously. Testers took the opportunity to select two major models at home and abroad - GPT-5.5, Fable-5, DeepSeek V4, Hunyuan 3 Preview - to test on the Beijing college entrance examination essay question "Planning and Effort", which requires no less than 700 Words can be used to write argumentative essays or narratives.
Each of the four major models wrote two argumentative essays and two narrative essays, which were cross-scored by four "evaluators". The result is as follows:
- DeepSeek V4 Narrative: Average score of 46 points (highest)
- Hunyuan 3 Preview: Hunyuan gave DeepSeek V4's narrative a perfect score of 60 points - "One's own people judge one's own people" triggered ridicule
- Each model is good at different styles of writing: GPT-5.5 and Fable-5 argumentative essays are more standardized, while DeepSeek V4 narrative essays are more literary.
The test also revealed the common characteristics of AI writing: neat structure and strong ability to quote classics, but lacks the "warmth" of real personal experience. This is especially obvious in the narrative scene - the AI writes stories smoothly but lacks real emotional depth.
AI writing college entrance examination essays is a fixed "private evaluation" every year, but its essential significance lies in revealing the differences in abilities of each model in Chinese creative writing. DeepSeek V4's advantage in narrative text shows that domestic models are catching up with international competitors in terms of Chinese literary expression, while Hunyuan's "full score" episode exposed the fundamental flaw of the AI self-evaluation system - the model lacks objective reference for scoring itself, and the construction of AI scoring standardization still has a long way to go.
The follow-up is worthy of attention:
- Objectivity of AI essay scoring: How to establish a cross-model standardized Chinese writing scoring system?
- Continuous evolution of Chinese creative writing: Can the literary advantages of DeepSeek V4 be maintained in subsequent versions?
- Standardized Application in Educational Scenarios: Boundaries and Standards of AI Writing in Educational Assessment and Assisted Teaching
Reviews