11 Mar 2026 - Deployment of the GPT-5.4 model
On their side, as of 4 Feb 2026 OpenAI increased the thinking effort for the Medium setting of the previously used GPT-5.2 model so much that the average cost rose by almost 30 % per message.
The point is that input tokens come at a certain price, whereas output tokens, which also include thinking tokens, are roughly 12x more expensive. Even a relatively small increase in reasoning effort therefore leads to a noticeable rise in cost.
On 5 Mar 2026 the GPT-5.4 model was released, which, according to both our internal benchmarks and, for example, ARC AGI v2, shows an increase in success rate while reducing cost. Its low effort looks better in quality than Medium on GPT-5.2.
On 10 Mar we therefore rolled out the new model into production with the default thinking effort setting set to Low. At the same time, our thinking moved toward giving advanced users more options and freedom to change the model's parameters - to get more thoughtful answers where it matters, or faster and cheaper ones if they don't insist on deeper exploration or don't mind waiting a bit longer. We therefore set the default thinking effort to Low with the option to configure it. It stayed like this in production all day. After collecting data, on 11 Mar 2026 we reverted the default thinking effort setting to Medium.
First production data
For comparison we present four consecutive days:
| Date | Model and setting | Positive out of total | Negative feedback | Average consumption |
|---|---|---|---|---|
| 9 Mar 2026 | GPT-5.2, Medium | 76 of 88 (86.4 %) | 12 | 8,316 credits |
| 10 Mar 2026 | GPT-5.4, low reasoning, low verbosity | 98 of 106 (92.5 %) | 8 | 3,337 credits |
| 11 Mar 2026 | GPT-5.4, medium reasoning, low verbosity | 108 of 113 (95.6 %) | 5 | 4,718 credits |
| 12 Mar 2026 | GPT-5.4, medium reasoning, low verbosity | 107 of 120 (89.2 %) | 13 | 4,422 credits |
The trend is shown graphically below:
The first day after deploying GPT-5.4 thus brought both more positive feedback and significantly lower average credit consumption compared to 9 Mar 2026. After reverting to Medium on 11 Mar 2026, the average consumption rose again, but remained below the level of the original GPT-5.2 setting.
We will continue to evaluate this and will gradually add recommendations for individual user groups and query types to the documentation.
