Skip to content
NEW SELECTThe new Itzi Select plans are here

Find the model that stays in character.

Human taste ranks the writing. Local tests expose memory, speed and the price you actually pay.

75% less

Save up to $7.11

Official $9.48 vs. $2.37 on Itzi.

6.8s

Fastest median

GPT-5.6 Luna

Gemini 3.7 Flash Thinking

Best value

84.0 RP quality at $0.052 per run.

The quality ranking

Quality blends 75% human preference with 25% local robustness.

Claude Fable 5

Best quality7/8 provisional

Anthropic

Missing episode penalized in robustness

91.6

/100

95% CI 87.9-94.9

Human taste

95.6 at 75%

Robustness

79.5 at 25%

Median

38.7s

P95

113.2s

Itzi estimate / run

Official est. $9.48

02

Claude Opus 4.7

Anthropic

90.5

/100

95% CI 87.6-93.4

Human taste

89.9 at 75%

Robustness

92.3 at 25%

Median

47.5s

P95

129.4s

Itzi estimate / run

Official est. $4.73

03

Claude Opus 4.6

Anthropic

89.1

/100

95% CI 86.1-91.9

Human taste

94.8 at 75%

Robustness

71.8 at 25%

Median

19.1s

P95

54.2s

Itzi estimate / run

Official est. $1.49

04

Gemini 3.7 Flash Thinking

Best value

Google

Provisional Arena sample

84.0

/100

95% CI 76.5-91.5

Human taste

86.0 at 75%

Robustness

77.8 at 25%

Median

47.9s

P95

204.2s

Itzi estimate / run

Official est. $0.21

05

GPT-5.6 Sol

OpenAI

83.6

/100

95% CI 79.5-87.6

Human taste

79.8 at 75%

Robustness

94.9 at 25%

Median

10s

P95

21.1s

Itzi estimate / run

Official est. $4.08

06

Claude Opus 4.8

Anthropic

83.3

/100

95% CI 80.4-86.2

Human taste

77.7 at 75%

Robustness

100.0 at 25%

Median

23.7s

P95

101.7s

Itzi estimate / run

Official est. $3.78

07

GLM 5.3

Z.ai

83.1

/100

95% CI 76.6-89.4

Human taste

79.1 at 75%

Robustness

94.9 at 25%

Median

70.2s

P95

118.4s

Itzi estimate / run

Official est. $1.17

08

Gemini 3.1 Pro Preview

Google

82.8

/100

95% CI 80.4-85.2

Human taste

81.3 at 75%

Robustness

87.2 at 25%

Median

17.1s

P95

23.3s

Itzi estimate / run

Official est. $1.28

09

Claude Opus 5

Anthropic

81.6

/100

95% CI 77.0-86.2

Human taste

78.1 at 75%

Robustness

92.3 at 25%

Median

26.9s

P95

132.4s

Itzi estimate / run

Official est. $3.01

10

Gemini 3.6 Flash Thinking

Google

81.0

/100

95% CI 77.0-85.0

Human taste

76.4 at 75%

Robustness

94.9 at 25%

Median

13s

P95

37.7s

Itzi estimate / run

Official est. $0.57

11

GPT-5.5

OpenAI

77.9

/100

95% CI 75.0-80.7

Human taste

72.2 at 75%

Robustness

94.9 at 25%

Median

20.1s

P95

42.1s

Itzi estimate / run

Official est. $2.41

12

Kimi K3

Moonshot AI

77.8

/100

95% CI 73.3-82.3

Human taste

77.2 at 75%

Robustness

79.5 at 25%

Median

9.9s

P95

19.4s

Itzi estimate / run

Official est. $0.79

13

Claude Sonnet 4.6

Anthropic

74.5

/100

95% CI 71.6-77.3

Human taste

70.2 at 75%

Robustness

87.2 at 25%

Median

13.2s

P95

50s

Itzi estimate / run

Official est. $0.70

14

GLM 5.1

Z.ai

72.5

/100

95% CI 69.5-75.4

Human taste

67.6 at 75%

Robustness

87.2 at 25%

Median

10.6s

P95

16.9s

Itzi estimate / run

Official est. $0.26

15

GLM 5.3 Flash

Z.ai

Provisional Arena sample

71.6

/100

95% CI 63.3-79.8

Human taste

63.8 at 75%

Robustness

94.9 at 25%

Median

92.9s

P95

197.6s

Itzi estimate / run

Official est. $0.16

16

Claude Sonnet 5

Anthropic

69.6

/100

95% CI 66.0-73.2

Human taste

63.7 at 75%

Robustness

87.2 at 25%

Median

48.2s

P95

143.3s

Itzi estimate / run

Official est. $3.40

17

GPT-5.6 Terra

OpenAI

68.0

/100

95% CI 64.2-71.8

Human taste

59.9 at 75%

Robustness

92.3 at 25%

Median

8.8s

P95

15.8s

Itzi estimate / run

Official est. $1.50

18

GLM 5.2

Z.ai

67.0

/100

95% CI 63.7-70.4

Human taste

68.0 at 75%

Robustness

64.1 at 25%

Median

10.3s

P95

24.8s

Itzi estimate / run

Official est. $0.28

19

DeepSeek V4 Pro

DeepSeek

Provisional Arena sample

64.7

/100

95% CI 57.1-72.3

Human taste

60.6 at 75%

Robustness

76.9 at 25%

Median

21.9s

P95

39.3s

Itzi estimate / run

Official est. $0.11

20

GPT-5.6 Luna

Fastest

OpenAI

58.6

/100

95% CI 54.5-62.6

Human taste

51.6 at 75%

Robustness

79.5 at 25%

Median

6.8s

P95

10.6s

Itzi estimate / run

Official est. $0.15

21

Grok 4.5

xAI

54.8

/100

95% CI 51.0-58.6

Human taste

68.8 at 75%

Robustness

12.8 at 25%

Median

56.8s

P95

181.4s

Itzi estimate / run

Official est. $0.10

22

DeepSeek V4 Flash

DeepSeek

54.8

/100

95% CI 51.6-57.9

Human taste

47.4 at 75%

Robustness

76.9 at 25%

Median

7.1s

P95

15s

Itzi estimate / run

Official est. $0.053

23

Grok 4.6

xAI

Provisional Arena sample

54.6

/100

95% CI 45.0-64.2

Human taste

64.3 at 75%

Robustness

25.6 at 25%

Median

24.2s

P95

48.3s

Itzi estimate / run

Official est. $0.22

Quality score vs. price

Higher RP quality moves up. Lower prices move right. The quality floor falls as models get cheaper, from 100 at $1.00 to 45 at $0.01.

At $1.00
Quality 100

Required RP quality

At $0.01
Quality 45

Required RP quality

Recommended

Best value: Gemini 3.7 Flash Thinking

Recommended

RP quality
89.1
Itzi run
$0.37

Best value

RP quality
84.0
Itzi run
$0.052

Recommended

RP quality
81.0
Itzi run
$0.14

Recommended

RP quality
72.5
Itzi run
$0.066

Recommended

RP quality
71.6
Itzi run
$0.041

Recommended

RP quality
64.7
Itzi run
$0.028

Recommended

RP quality
54.8
Itzi run
$0.013

Outside the zone

Claude Fable 5

7/8 provisional

Misses RP quality and cost targets

RP quality
91.6
Itzi run
$2.37

Claude Opus 4.7

Misses RP quality and cost targets

RP quality
90.5
Itzi run
$1.18

GPT-5.6 Sol

Misses RP quality and cost targets

RP quality
83.6
Itzi run
$1.02

Claude Opus 4.8

Misses RP quality target

RP quality
83.3
Itzi run
$0.94

GLM 5.3

Misses RP quality target

RP quality
83.1
Itzi run
$0.29

Gemini 3.1 Pro Preview

Misses RP quality target

RP quality
82.8
Itzi run
$0.32

Claude Opus 5

Misses RP quality target

RP quality
81.6
Itzi run
$0.75

GPT-5.5

Misses RP quality target

RP quality
77.9
Itzi run
$0.60

Kimi K3

Misses RP quality target

RP quality
77.8
Itzi run
$0.20

Claude Sonnet 4.6

Misses RP quality target

RP quality
74.5
Itzi run
$0.17

Claude Sonnet 5

Misses RP quality target

RP quality
69.6
Itzi run
$0.85

GPT-5.6 Terra

Misses RP quality target

RP quality
68.0
Itzi run
$0.38

GLM 5.2

Misses RP quality target

RP quality
67.0
Itzi run
$0.069

GPT-5.6 Luna

Misses RP quality target

RP quality
58.6
Itzi run
$0.039

Grok 4.5

Misses RP quality target

RP quality
54.8
Itzi run
$0.025

Grok 4.6

Misses RP quality target

RP quality
54.6
Itzi run
$0.055

Rose marks models above the price-adjusted quality floor.

A dashed ring marks a 7/8 provisional result. Exact values remain beside every model.

RP quality and price chart source values
ModelRP qualityEstimated Itzi benchmark-run costRequired RP quality at this priceRecommendedStatus
Claude Fable 591.6$2.37100.0No7 of 8 provisional
Claude Opus 4.790.5$1.18100.0NoComplete
Claude Opus 4.689.1$0.3788.2YesComplete
Gemini 3.7 Flash Thinking84.0$0.05264.7YesComplete
GPT-5.6 Sol83.6$1.02100.0NoComplete
Claude Opus 4.883.3$0.9499.3NoComplete
GLM 5.383.1$0.2985.3NoComplete
Gemini 3.1 Pro Preview82.8$0.3286.4NoComplete
Claude Opus 581.6$0.7596.6NoComplete
Gemini 3.6 Flash Thinking81.0$0.1476.8YesComplete
GPT-5.577.9$0.6094.0NoComplete
Kimi K377.8$0.2080.6NoComplete
Claude Sonnet 4.674.5$0.1779.2NoComplete
GLM 5.172.5$0.06667.5YesComplete
GLM 5.3 Flash71.6$0.04161.9YesComplete
Claude Sonnet 569.6$0.8598.0NoComplete
GPT-5.6 Terra68.0$0.3888.3NoComplete
GLM 5.267.0$0.06968.1NoComplete
DeepSeek V4 Pro64.7$0.02857.4YesComplete
GPT-5.6 Luna58.6$0.03961.1NoComplete
Grok 4.554.8$0.02556.0NoComplete
DeepSeek V4 Flash54.8$0.01348.3YesComplete
Grok 4.654.6$0.05565.4NoComplete

With Itzi, you save

Up to $7.11 in one run.

Compared with the official PAYG estimate for the same complete benchmark run.

What one benchmark run means

8 episodes and 64 model responses at maximum reasoning. Individual usage varies.

Estimated benchmark-run cost

Average estimate from observed token and cache usage. Actual spend varies with prompt length, response length, reasoning and cache usage.

Claude Fable 5

7/8 provisional
Official estimate
$9.48
On Itzi
You save
Official
$9.48
Itzi
$2.37

Claude Opus 4.7

Official estimate
$4.73
On Itzi
You save
Official
$4.73
Itzi
$1.18

Claude Opus 4.6

Official estimate
$1.49
On Itzi
You save
Official
$1.49
Itzi
$0.37

Gemini 3.7 Flash Thinking

Official estimate
$0.21
On Itzi
You save
Official
$0.21
Itzi
$0.052

GPT-5.6 Sol

Official estimate
$4.08
On Itzi
You save
Official
$4.08
Itzi
$1.02

Claude Opus 4.8

Official estimate
$3.78
On Itzi
You save
Official
$3.78
Itzi
$0.94

GLM 5.3

Official estimate
$1.17
On Itzi
You save
Official
$1.17
Itzi
$0.29

Gemini 3.1 Pro Preview

Official estimate
$1.28
On Itzi
You save
Official
$1.28
Itzi
$0.32

Claude Opus 5

Official estimate
$3.01
On Itzi
You save
Official
$3.01
Itzi
$0.75

Gemini 3.6 Flash Thinking

Official estimate
$0.57
On Itzi
You save
Official
$0.57
Itzi
$0.14

GPT-5.5

Official estimate
$2.41
On Itzi
You save
Official
$2.41
Itzi
$0.60

Kimi K3

Official estimate
$0.79
On Itzi
You save
Official
$0.79
Itzi
$0.20

Claude Sonnet 4.6

Official estimate
$0.70
On Itzi
You save
Official
$0.70
Itzi
$0.17

GLM 5.1

Official estimate
$0.26
On Itzi
You save
Official
$0.26
Itzi
$0.066

GLM 5.3 Flash

Official estimate
$0.16
On Itzi
You save
Official
$0.16
Itzi
$0.041

Claude Sonnet 5

Official estimate
$3.40
On Itzi
You save
Official
$3.40
Itzi
$0.85

GPT-5.6 Terra

Official estimate
$1.50
On Itzi
You save
Official
$1.50
Itzi
$0.38

GLM 5.2

Official estimate
$0.28
On Itzi
You save
Official
$0.28
Itzi
$0.069

DeepSeek V4 Pro

Official estimate
$0.11
On Itzi
You save
Official
$0.11
Itzi
$0.028

GPT-5.6 Luna

Official estimate
$0.15
On Itzi
You save
Official
$0.15
Itzi
$0.039

Grok 4.5

Official estimate
$0.10
On Itzi
You save
Official
$0.10
Itzi
$0.025

DeepSeek V4 Flash

Official estimate
$0.053
On Itzi
You save
Official
$0.053
Itzi
$0.013

Grok 4.6

Official estimate
$0.22
On Itzi
You save
Official
$0.22
Itzi
$0.055

How the benchmark works

See how quality, speed and benchmark-run prices are calculated.

Read the methodology

Human taste

Arena contributes 75% of the score: 50% Creative Writing, 20% Multi-Turn, 20% Hard Prompts English and 10% Overall from a snapshot captured across 1-2 Sep 2026.

Local robustness

Eight scripted episodes audit character fidelity, continuity, agency, spatial causality and instruction retention. A 7/8 result is provisional, and every objective criterion from the missing episode counts as a failure.

Speed

Median and p95 latency come from completed benchmark turns.

Public PAYG price

Benchmark-run prices apply observed token and cache usage to 64 model responses at public PAYG rates. Actual spend varies. The comparison excludes internal routing costs.

Fable 5 is ranked provisionally from 7 of 8 completed episodes. Its missing episode counts as failed robustness criteria. GLM 5.3, GLM 5.3 Flash, DeepSeek V4 Pro and Grok 4.6 use provisional Arena samples.