The development of AI models by 2026 shows a clear direction: competition is no longer just about who has the highest benchmark scores. The latest models are also racing to become faster, more token-efficient, capable of completing longer tasks, and easier to use in everyday workflows.
This change is significant because the costs of AI do not only arise from subscription prices. There is time spent reviewing results, costs associated with repeating instructions, risks of errors, and additional work when a model produces something that seems convincing but is ultimately unusable.
From the Smartest Models to the Most Completed Work
On July 9, 2026, OpenAI introduced the GPT-5.6 family, emphasizing performance per dollar, more efficient token usage, and the ability to complete longer knowledge and coding tasks. OpenAI also stated that GPT-5.6 Sol is designed to reduce the number of work iterations and the time required for tool-based tasks. ([openai.com](https://openai.com/index/gpt-5-6/?utm_source=openai))
Anthropic took a similar direction when introducing Claude Opus 5 on July 24, 2026. The company highlighted improvements in efficiency, capabilities in knowledge work, and better outcomes at a claimed lower cost compared to previous models. ([anthropic.com](https://www.anthropic.com/news/claude-opus-5?src_trk=em6714483831e7b6.91161670910169237&utm_source=openai))
These figures still need to be interpreted carefully. Benchmarks are test results under specific conditions, while real work often involves messy data, changing instructions, limited access, and the need to coordinate with other humans. High scores can signal capability, but they do not guarantee that a model will complete our tasks correctly.
Cheap Tokens Do Not Necessarily Mean Cheap Work
Tokens are units of text processed by AI models. Many service providers use the number of tokens to calculate usage costs. Therefore, a model that generates long answers is not necessarily more expensive directly if its price is low. However, the actual costs are broader than just tokens.
- Review Costs: AI results still need to be read, compared with sources, and corrected.
- Repetition Costs: Unclear instructions can lead users to repeat processes multiple times.
- Error Costs: Mistakes in reports, code, or analysis can lead to far more expensive consequences.
- Coordination Costs: AI results may not fit the existing formats, systems, or approval processes.
Therefore, a more useful metric is cost per usable output. For example, it is not just about calculating how many cents it costs to create a summary, but how long it takes until that summary can be delivered without major revisions.
Why Long-Form Capability Is Becoming Important
Next-generation AI models are increasingly directed to handle tasks that require multiple stages: understanding documents, searching for information, drafting, using tools, and then reviewing the results. OpenAI reports that the use of Codex is shifting from short requests to tasks expected to require more than one hour of human work. ([openai.com](https://openai.com/index/how-agents-are-transforming-work/?utm_source=openai))
This marks a shift from AI as a “answer machine” to AI as a “work machine.” The difference is akin to asking someone to answer a single question versus asking them to prepare a complete report. In the latter task, the ability to plan, maintain context, recognize deadlocks, and request clarification becomes much more important.
However, the longer the delegated task, the greater the potential for errors. Models can make incorrect assumptions in the early stages and then build the entire output based on those assumptions. A small mistake at the first step can turn into a significant problem in the final result.
What This Means for Users and Small Teams
For the average user, the development of AI models does not mean we must always subscribe to the most expensive models. A more sensible choice is to match the model's capabilities with the risk level of the work.
- Use fast and cheap models for tidying up notes, creating title variations, or changing text formats.
- Use more powerful models for analysis involving multiple documents, debugging, or tasks requiring step-by-step reasoning.
- For important decisions, use AI as a material compiler, not the final decision-maker.
- If the work involves personal data or company information, check data storage and usage policies before uploading.
In practice, the “best” model may differ for each team. The top model for coding may not be the best for writing legal documents. A fast model for drafting may not be suitable for analyzing spreadsheets with many exceptions.
How to Measure AI Without Getting Trapped in Demos
Small teams can conduct simple tests using 10 to 20 examples of real work. Examples include customer emails, weekly reports, database queries, contract documents, or frequently repeated administrative tasks.
- Define the final output. Explain what is meant by completed work, not just a good answer.
- Record total time. Calculate the time spent running AI, reviewing results, and making corrections.
- Measure error rates. Note factual, formatting, logical errors, and missed information.
- Compare costs. Include subscription costs, API usage, tool usage, and human time.
- Test consistency. Run several similar examples to see if the results are stable.
A simple formula could be: AI value = time saved minus review time and error costs. This is not an official accounting formula, but it is helpful to ensure evaluations do not stop at the impression that AI results look impressive.
Conclusion: Efficiency Must Be Proven in Workflows
The latest AI models are indeed moving towards being more powerful, longer-lasting, and more efficient. However, increased capability does not automatically lead to productivity. Productivity only emerges when models can integrate into workflows without creating too much additional work.
Sources & Further Reading
- GPT-5.6: Frontier intelligence that scales with your ambition
- Claude Opus 5
- How agents are transforming work
– Rio Yotto @rioyotto
