OpenAI's GPT-6 Astra Achieves First Perfect Score on Korean CSAT Using Record-Low Tokens
Key Takeaways
- •OpenAI's GPT-6 Astra scored a perfect 450 in the 2026 CSAT LLM Solution Log, the first model to do so across all tested subjects.
- •Astra consumed 357,000 tokens, the lowest reported total among publicly available models, which lowers inference costs relative to rivals.
- •OpenAI's GPT-5.6 ranked second with 448.5 points, ahead of GPT-5.4 at 448, Claude Fable 5.1 at 447.5, and Gemini 3.1 Pro at 445.
- •The model's efficiency is credited to a design that retains, compresses, and reuses context from earlier reasoning steps during repetitive tasks.
- •Industry experts caution that AI capability should be judged on open-ended, real-world problem-solving rather than standardized exam results alone.

OpenAI's latest artificial intelligence model has earned a perfect score on South Korea's College Scholastic Ability Test (CSAT) for the first time, according to a newly published analysis. While top-tier AI models have previously posted near-perfect results, the outcome is drawing particular attention because the model achieved it while consuming fewer tokens than its rivals — a metric that matters commercially, since token usage translates directly into inference costs for developers and enterprises running these models at scale.
Results posted on GitHub on Sunday showed that OpenAI's new GPT-6 Astra was the only evaluated model to score a perfect 450 in the 2026 CSAT LLM Solution Log. The evaluation tested models on questions from the 2026 CSAT covering Korean language, English, mathematics, Korean history and four elective subjects — Physics I, Chemistry I, Life Science I and Society and Culture — without internet access. The CSAT, a standardized exam taken by hundreds of thousands of Korean students each year, has become a recurring yardstick for comparing how well frontier models handle multilingual, multi-subject reasoning under exam conditions.
Astra was the first AI model to achieve a perfect score across all tested subjects in the evaluation. In February, Google's Gemini 3.1 Pro earned a perfect score while taking only two electives — Chemistry I and Life Science I — but missed questions in Physics I and a social studies subject under the broader format used in this evaluation.
In the rankings released Sunday, OpenAI's GPT-5.6 placed second with 448.5 points, followed by GPT-5.4 in third with 448. Anthropic's Claude Fable 5.1 placed fourth with 447.5, and Gemini 3.1 Pro came in fifth with 445.
GPT-6 Astra used 357,000 tokens, the lowest reported total among publicly available models. By comparison, GPT-5.6 used 429,000 tokens and Claude Fable 5.1 used 562,000. Because the exam questions and answers were publicly available, the result cannot rule out prior exposure through training data, though competitor models such as Claude Fable 5.1 and Gemini 3.1 Pro faced the same conditions.
The model's strong performance combined with lower token usage is attributed to a reasoning-linked design that retains context during repetitive tasks. A provider adapter harness compresses and reuses information from earlier reasoning steps, allowing the AI to take the test in a manner that mimics the approach of a top-scoring student.
Lee Seung-hyun, an adjunct professor at Hanyang University, said the key is not simply that the model became smarter, but that it can retain prior reasoning steps, compress context and revise its plan. "It means that rather than thinking harder each time, it continued with what it had already figured out before," Lee said.
OpenAI described the result as encouraging, noting that the model achieved efficient processing using hyperscale AI computing infrastructure. An OpenAI official said the milestone demonstrates that models can become more powerful while simultaneously improving performance and lowering costs.
Industry experts cautioned, however, that a model's potential should be judged by its ability to solve open-ended problems rather than by standardized exam scores alone. A Korean AI industry insider said it will be more meaningful if AI can tackle real-world workplace challenges where there are no predetermined answers. How leading models perform on broader independent benchmarks — and whether rivals close the efficiency gap — will show whether this result marks a durable shift in the cost-performance trade-off.
This article is based on a report by the Hankook Ilbo, the sister publication of The Korea Times, translated by a generative AI system and edited by The Korea Times. Original source: The Korea Times.