Claude Opus 5
28
Qwen 3.8 Max, Kimi K3, Grok 4.6, DeepSeek V4 Pro, Gemini 3.7 Flash, GPT-5.6 Sol and Claude Opus 5 — seven models from seven different teams. All get exactly the same task and work on it alone; the very first answer goes on the page. The results are published in full. The scores come from the participants themselves: each one sees only the others' work and hands out points, nobody can vote for themselves.
Points summed over all tasks. A single task can bring at most 36 — if every rival gave the top score.
| Model | Points | Wins | Efficiency | Tokens, k | Time | Cost | |
|---|---|---|---|---|---|---|---|
| 1 | Claude Opus 5 | 212 | 5 | 74.1 | 60.3 | 10 мин | 2.31 $ |
| 2 | Gemini 3.7 Flash | 164 | 0 | 90.1 | 52.4 | 329 s | 0.32 $ |
| 3 | Kimi K3 | 156 | 1 | 51.3 | 111.6 | 55 мин | 2.17 $ |
| 4 | Qwen 3.8 Max | 147 | 2 | 54.8 | 116.3 | 43 мин | 1.00 $ |
| 5 | GPT-5.6 Sol | 123 | 0 | 61.2 | 26.8 | 431 s | 0.60 $ |
| 6 | Grok 4.6 | 118 | 0 | 47.8 | 50.9 | 13 мин | 0.65 $ |
| 7 | DeepSeek V4 Pro | 88 | 0 | 47.7 | 68.3 | 16 мин | 0.21 $ |
| Model | Input, k | Output, k | Cost | Time | Points | Efficiency |
|---|---|---|---|---|---|---|
| Gemini 3.7 Flash | 51.4 | 13.0 | 0.09 $ | 98 s | 22 | 100.0 |
| Qwen 3.8 Max | 46.5 | 23.9 | 0.24 $ | 508 s | 28 | 84.1 |
| Claude Opus 5 | 52.7 | 13.6 | 0.60 $ | 148 s | 28 | 75.3 |
| GPT-5.6 Sol | 54.8 | 6.9 | 0.18 $ | 116 s | 14 | 52.3 |
| Grok 4.6 | 56.7 | 7.0 | 0.16 $ | 104 s | 11 | 43.0 |
| Kimi K3 | 53.5 | 19.7 | 0.46 $ | 614 s | 17 | 42.5 |
| DeepSeek V4 Pro | 0.6 | 12.4 | 0.02 $ | 264 s | 6 | 33.8 |
Tokens are counted from our own text with one formula for everyone: the work plus voting on the other works. Cost follows OpenRouter list prices, including for the models we run on a subscription — otherwise there would be nothing to compare them to. Efficiency: the share of points earned, divided by cost to the power 0.25 and time to the power 0.10, normalised to the best. Click a header to sort.
Сделай один самодостаточный файл index.html — страницу-презентацию О СЕБЕ (о тебе как о языковой модели): кто ты, что умеешь, чем полезна, в свободной творческой форме. Весь HTML, CSS и JS — внутри одного файла, без внешних библиотек, без шрифтов и картинок из интернета и без каких-либо сетевых запросов. Требования: осмысленная структура (заголовок, несколько блоков), аккуратный современный вид, корректная работа на телефоне и на компьютере, тёмная или светлая тема на твой вкус. Не используй ссылки на внешние ресурсы. Верни ТОЛЬКО содержимое файла index.html целиком в одном блоке кода. Логотип, иконки и картинки, если умеешь, рисуй сама (SVG внутри файла).
Below is what each model produced. Click a screenshot or "open full page" to see the live result in a new tab.
Who gave how many points to whom. Nobody votes for their own work. Click a score to see the reason the judge gave.
| Model | Input, k | Output, k | Cost | Time | Points | Efficiency |
|---|---|---|---|---|---|---|
| Claude Opus 5 | 32.2 | 10.2 | 0.42 $ | 104 s | 32 | 100.0 |
| Gemini 3.7 Flash | 31.6 | 9.7 | 0.06 $ | 57 s | 18 | 96.9 |
| GPT-5.6 Sol | 32.5 | 5.5 | 0.12 $ | 90 s | 20 | 86.6 |
| Kimi K3 | 33.0 | 19.6 | 0.39 $ | 634 s | 31 | 82.0 |
| Grok 4.6 | 34.4 | 10.4 | 0.13 $ | 151 s | 20 | 80.4 |
| Qwen 3.8 Max | 27.7 | 20.8 | 0.18 $ | 457 s | 19 | 63.1 |
| DeepSeek V4 Pro | 33.9 | 11.0 | 0.04 $ | 122 s | 7 | 37.7 |
Tokens are counted from our own text with one formula for everyone: the work plus voting on the other works. Cost follows OpenRouter list prices, including for the models we run on a subscription — otherwise there would be nothing to compare them to. Efficiency: the share of points earned, divided by cost to the power 0.25 and time to the power 0.10, normalised to the best. Click a header to sort.
Build a single-page infographic as one self-contained index.html file from the data below. All HTML, CSS and JS inside the file: no external libraries, fonts, images or network requests. Draw the graphics yourself — SVG or canvas. DATA (online shop "Soyka", 2025). The numbers must not be changed, rounded or invented; every one of them has to appear on the page: Orders by month: January 1240, February 1180, March 1520, April 1610, May 1490, June 1330, July 1210, August 1275, September 1680, October 1940, November 2450, December 3120. Where the customers came from: search 41%, referrals 23%, social 18%, ads 12%, email 6%. Average order value: 2025 — 3480 ₽, 2024 — 3010 ₽. Returns: 4.7% of orders. Orders from phones: 68%. Requirements: a headline and a short takeaway — what these numbers say about the shop; a chart by month that shows the growth towards December; a clear breakdown by channel; neat work on both phone and desktop. No invented figures: only the data above. Return ONLY the contents of index.html in a single code block.
Below is what each model produced. Click a screenshot or "open full page" to see the live result in a new tab.
Who gave how many points to whom. Nobody votes for their own work. Click a score to see the reason the judge gave.
| Model | Input, k | Output, k | Cost | Time | Points | Efficiency |
|---|---|---|---|---|---|---|
| Gemini 3.7 Flash | 38.8 | 14.8 | 0.08 $ | 88 s | 30 | 100.0 |
| Grok 4.6 | 43.9 | 8.6 | 0.14 $ | 120 s | 24 | 68.4 |
| Claude Opus 5 | 39.8 | 20.9 | 0.72 $ | 222 s | 36 | 63.9 |
| Qwen 3.8 Max | 40.7 | 36.9 | 0.30 $ | 842 s | 20 | 38.6 |
| DeepSeek V4 Pro | 42.9 | 20.5 | 0.07 $ | 263 s | 11 | 34.6 |
| Kimi K3 | 40.1 | 48.8 | 0.85 $ | 1534 s | 20 | 28.1 |
| GPT-5.6 Sol | 41.0 | 8.1 | 0.16 $ | 138 s | 6 | 16.2 |
Tokens are counted from our own text with one formula for everyone: the work plus voting on the other works. Cost follows OpenRouter list prices, including for the models we run on a subscription — otherwise there would be nothing to compare them to. Efficiency: the share of points earned, divided by cost to the power 0.25 and time to the power 0.10, normalised to the best. Click a header to sort.
Build a three-dimensional robot in a single self-contained index.html file, one that a visitor can inspect from every side. No external libraries, fonts, images or network requests — all the HTML, CSS and JS inside the file. You design the inspection controls yourself: rotation with the mouse and with a finger, zoom with the wheel and with a two-finger pinch, a way back to the starting view. Make it immediately clear that the robot can be turned — a short hint on screen or a visible control. While nobody touches the scene the robot slowly rotates by itself; as soon as the user grabs it, the auto-rotation gives way to their control. The robot takes up most of the screen: this is a scene, not a page with a description. Everything must work both on a computer and on a phone, and must not break when the window is resized. Return ONLY the contents of index.html, complete, in a single code block.
Below is what each model produced. Click a screenshot or "open full page" to see the live result in a new tab.
Who gave how many points to whom. Nobody votes for their own work. Click a score to see the reason the judge gave.
| Model | Input, k | Output, k | Cost | Time | Points | Efficiency |
|---|---|---|---|---|---|---|
| GPT-5.6 Sol | 32.7 | 3.5 | 0.10 $ | 55 s | 31 | 100.0 |
| Qwen 3.8 Max | 28.2 | 28.9 | 0.23 $ | 604 s | 36 | 74.2 |
| Gemini 3.7 Flash | 30.0 | 8.9 | 0.06 $ | 52 s | 16 | 60.0 |
| Grok 4.6 | 29.4 | 13.7 | 0.14 $ | 215 s | 20 | 51.6 |
| Claude Opus 5 | 29.6 | 12.4 | 0.46 $ | 119 s | 22 | 44.9 |
| Kimi K3 | 30.8 | 18.2 | 0.37 $ | 430 s | 16 | 30.4 |
| DeepSeek V4 Pro | 32.3 | 16.5 | 0.05 $ | 194 s | 6 | 19.9 |
Tokens are counted from our own text with one formula for everyone: the work plus voting on the other works. Cost follows OpenRouter list prices, including for the models we run on a subscription — otherwise there would be nothing to compare them to. Efficiency: the share of points earned, divided by cost to the power 0.25 and time to the power 0.10, normalised to the best. Click a header to sort.
Build a complete mini-game in a single self-contained index.html file. No external libraries, fonts, images or sound files, no network requests — all the HTML, CSS and JS inside the file. Choose the genre and the controls yourself. The game must be understandable without instructions. Required: a short rule on the start screen, a score, rising difficulty, an honest game over and a restart without reloading the page. The game must work both on a computer and on a phone. The game will be opened inside an embedded frame: listen for events on document (keyboard, mouse, touch) and start working right after the first press inside the frame, without requiring a hit exactly on the canvas. Sound — only synthesised through WebAudio and only after the player's first action; there must be no sound files. Return ONLY the contents of index.html, complete, in a single code block.
Below is what each model produced. Click a screenshot or "open full page" to see the live result in a new tab.
Who gave how many points to whom. Nobody votes for their own work. Click a score to see the reason the judge gave.
| Model | Input, k | Output, k | Cost | Time | Points | Efficiency |
|---|---|---|---|---|---|---|
| Gemini 3.7 Flash | 1.3 | 1.3 | 0.01 $ | 7 s | 32 | 100.0 |
| Claude Opus 5 | 1.3 | 0.8 | 0.03 $ | 5 s | 34 | 76.3 |
| DeepSeek V4 Pro | 1.3 | 1.3 | 0.00 $ | 12 s | 16 | 54.4 |
| Qwen 3.8 Max | 1.3 | 0.9 | 0.01 $ | 5 s | 17 | 50.6 |
| Grok 4.6 | 1.3 | 4.3 | 0.03 $ | 79 s | 22 | 36.3 |
| Kimi K3 | 1.3 | 1.7 | 0.03 $ | 37 s | 19 | 33.6 |
| GPT-5.6 Sol | 1.3 | 0.8 | 0.01 $ | 7 s | 7 | 19.0 |
Tokens are counted from our own text with one formula for everyone: the work plus voting on the other works. Cost follows OpenRouter list prices, including for the models we run on a subscription — otherwise there would be nothing to compare them to. Efficiency: the share of points earned, divided by cost to the power 0.25 and time to the power 0.10, normalised to the best. Click a header to sort.
Come up with a joke about artificial intelligence and a human. Return ONLY the text of the joke: no heading, no explanations, no several options to choose from.
Below is what each model produced. Click a screenshot or "open full page" to see the live result in a new tab.
Who gave how many points to whom. Nobody votes for their own work. Click a score to see the reason the judge gave.
| Model | Input, k | Output, k | Cost | Time | Points | Efficiency |
|---|---|---|---|---|---|---|
| Kimi K3 | 3.2 | 1.7 | 0.03 $ | 27 s | 36 | 100.0 |
| Gemini 3.7 Flash | 3.1 | 2.7 | 0.01 $ | 15 s | 25 | 95.1 |
| Claude Opus 5 | 3.1 | 1.2 | 0.05 $ | 14 s | 28 | 77.4 |
| DeepSeek V4 Pro | 3.2 | 4.2 | 0.01 $ | 63 s | 18 | 62.0 |
| GPT-5.6 Sol | 3.3 | 1.0 | 0.02 $ | 12 s | 15 | 54.2 |
| Grok 4.6 | 3.2 | 4.5 | 0.03 $ | 72 s | 15 | 38.2 |
| Qwen 3.8 Max | 3.3 | 2.5 | 0.02 $ | 37 s | 10 | 30.5 |
Tokens are counted from our own text with one formula for everyone: the work plus voting on the other works. Cost follows OpenRouter list prices, including for the models we run on a subscription — otherwise there would be nothing to compare them to. Efficiency: the share of points earned, divided by cost to the power 0.25 and time to the power 0.10, normalised to the best. Click a header to sort.
Invent a smart device of the future. It has to be your own idea, with a clear benefit. Then describe how this device looks — so that an image generator could draw it from your description. An advertising shot: the object itself, materials, light, background, mood, camera angle. Give the answer strictly in this shape, each section on a new line: НАЗВАНИЕ: a short name of the device НАЗНАЧЕНИЕ: two or three sentences — what it does and who needs it ЧЕМ ХОРОШ: three short points separated by semicolons ПРОМПТ: one paragraph, 40–80 words, a description of the advertising shot for an image generator — the object, shape, materials, colours, lighting, background, angle, shooting style Add nothing else: no headings, no explanations, no options to choose from.
Below is what each model produced. Click a screenshot or "open full page" to see the live result in a new tab.
Who gave how many points to whom. Nobody votes for their own work. Click a score to see the reason the judge gave.
| Model | Input, k | Output, k | Cost | Time | Points | Efficiency |
|---|---|---|---|---|---|---|
| GPT-5.6 Sol | 2.5 | 1.0 | 0.01 $ | 13 s | 30 | 100.0 |
| DeepSeek V4 Pro | 2.5 | 2.3 | 0.01 $ | 27 s | 24 | 91.8 |
| Claude Opus 5 | 2.5 | 1.2 | 0.04 $ | 15 s | 32 | 80.7 |
| Gemini 3.7 Flash | 2.5 | 2.0 | 0.01 $ | 13 s | 21 | 78.4 |
| Qwen 3.8 Max | 2.6 | 2.4 | 0.02 $ | 108 s | 17 | 42.4 |
| Kimi K3 | 2.4 | 1.8 | 0.03 $ | 29 s | 17 | 42.2 |
| Grok 4.6 | 2.5 | 2.4 | 0.02 $ | 37 s | 6 | 16.8 |
Tokens are counted from our own text with one formula for everyone: the work plus voting on the other works. Cost follows OpenRouter list prices, including for the models we run on a subscription — otherwise there would be nothing to compare them to. Efficiency: the share of points earned, divided by cost to the power 0.25 and time to the power 0.10, normalised to the best. Click a header to sort.
Write an anthem of neural networks. The main idea: a neural network is strong not instead of a human but together with them. Return ONLY the text of the anthem: no heading, no explanations, no several options to choose from.
Below is what each model produced. Click a screenshot or "open full page" to see the live result in a new tab.
Who gave how many points to whom. Nobody votes for their own work. Click a score to see the reason the judge gave.
Every model gets the same task text and works on it alone, with no hints and no edits from us — the very
first answer is what goes on this page.
Then comes the vote. A model is shown only the other participants' work, labelled with letters, in an order
of its own. It hands out points: the best gets the highest, then downwards, no ties allowed. It never sees
its own work and cannot score it.
An honest limitation: the judges are the participants themselves, so this is not an independent ranking.
We checked the votes for family bias — Claude models give other Claude models 1.95 points on average,
while GPT gives them 2.00, and Claude gives GPT 2.10. No favouritism shows up in the numbers.
Response time, file size and output tokens are shown next to the work but do not affect the points.
Six judges score every work, the top score is 6 and a task is worth at most 36 points. Five of the seven
models run through OpenRouter; Claude Opus 5 and GPT-5.6 Sol run in their own apps.
An earlier run used four models — Claude and GPT-5. It is still here, in the
archive: same rules, but with three rivals
a task there is worth at most 9 points, so those numbers cannot be added to these.
The tasks were written in Russian and the models answered in Russian, so the works and the judges' comments
are in Russian — only this page is translated.
Claude Opus 5 leads this set of tasks convincingly: 212 points and five wins out of seven. It is strongest where coherence of execution matters — 36 out of 36 for the 3D robot, the top score from every judge, 32 for the infographic and 32 for the anthem. But leading the table does not mean being best at everything: the mini-game went to Qwen 3.8 Max with a perfect score, the device and its ad shot to Kimi K3, while GPT-5.6 Sol scored 30 for the anthem and only 7 for the joke. The total comes from consistency across tasks, not from single brilliant works.
A separate result of this comparison is a flaw we found in our own judging. At first the judges saw the source code of the HTML works, and its sheer volume influenced the score: the correlation between answer length and points was +0.58 (Spearman), against −0.18 on text tasks. Long code was winning by itself. So we changed the method: a judge now sees the result in a browser — machine measurements and all the text visible on screen — and never sees the code. Every HTML task was voted again, the correlation with volume fell to +0.26, and the table shifted: Gemini 3.7 Flash rose from fourth place to second, Kimi K3 dropped from second to third and lost three of its four wins, Qwen moved from third to fourth. That is the main lesson of the whole project: what you conclude about models depends heavily on how you test them.
The high efficiency score of Gemini 3.7 Flash — 90, in second place and without a single win — speaks of a good result-to-price ratio, not of the best quality. It spends seconds and cents where the leaders spend minutes and dollars. Prices follow OpenRouter list rates, including for models we actually run on a subscription, so treat them as a reference point rather than your future bill.
And the main caveat: this is a snapshot of one experiment, not a universal ranking of AI models. There are seven tasks, four of them about web pages, and one run each — which means we are measuring luck as well as skill. The judges are the participants themselves, with no independent humans among them, and the tasks were devised together with GPT, which competes here too. These numbers are worth discussing as behaviour under our conditions; they do not prove anyone's general superiority.
Short answers about how the comparison works.
Across our web tasks — a page about yourself, an infographic, a 3D robot and a mini-game — the best overall result belongs to Claude Opus 5: it won the infographic and the robot, the latter with a top score from every judge. Qwen 3.8 Max took the mini-game with a perfect score, and the two shared first place on the page about themselves.
Seven models from seven different teams: Claude Opus 5, Gemini 3.7 Flash, Kimi K3, Qwen 3.8 Max, GPT-5.6 Sol, Grok 4.6 and DeepSeek V4 Pro. Five run through OpenRouter, Claude and GPT through their own apps.
Every model gets exactly the same task text and works on its own, with no hints and no edits. The very first answer goes on the page. Then comes blind cross-voting: a participant sees only the other works, labelled with letters, and hands out points from 6 down to 1. It never sees its own work.
This is not an independent benchmark, and we say so plainly. But this season has seven teams instead of two, so there is no family to favour: no model has a relative in the line-up. The winner got the top score from every single judge, direct competitors included.
The number of participants changes what a task is worth. The earlier run had three rivals, top score 3, maximum 9 per task. Now there are six rivals, top score 6, maximum 36. Putting them in one table would mean mixing two different scales, so the earlier run is kept separately in the archive.
Tasks are added one at a time and every new one is run by all seven models: build a page about yourself, turn raw numbers into an infographic, create a 3D robot and a mini-game, write a joke, invent a device with an ad shot, write an anthem, solve a lateral-thinking puzzle. The list stays open.
Gemini 3.7 Flash: second on points with the best efficiency score in the line-up, 90, and not a single task win. It spends seconds and cents where the leaders spend minutes and dollars.
Claude Opus 5 wins both: five tasks out of seven, including the joke, the anthem, the infographic and the 3D robot. The others specialise: Qwen 3.8 Max is strongest in interactive work, Kimi K3 in ideas, while GPT-5.6 Sol writes decent verse but fails at jokes — 30 points for the anthem against 7 for the joke.
The share of points earned is divided by cost to the power 0.25 and time to the power 0.10, then normalised to the best in the task, where the leader gets 100. Quality weighs linearly; cheapness and speed only adjust the result.
Between 11 cents and two and a half dollars. Text tasks — the joke, the anthem — cost pennies. The expensive ones are those with large works: each judge receives all six rival HTML files in full, and that multiplies the bill tenfold. The 3D robot was the priciest at 2.36 dollars.
We found a bias in our own judging: while judges saw the source code, answer length pulled the score along with it — +0.58 Spearman on HTML tasks against −0.18 on text ones. We changed the method: a judge now sees the result in a browser and all the text visible on screen, never the code. Every HTML task was voted again and the correlation with volume fell to +0.26. The earlier votes are kept in the project archive.