The difference between chatbots and AI agents is becoming clear. Chatbots typically answer questions and then stop. AI agents can receive goals, break them down into several steps, use tools, read files, run code, and repeat the process until the job is done.
This change is significant because the way we work with AI is also changing. We are no longer just asking for a "summary," but can assign tasks such as "check the causes of the API error increase, compare it with the last deployment changes, and then save the findings and recommendations."
On September 10, 2026, OpenAI introduced the Agents API in public beta. This platform is designed to run agents in long sessions, using a working environment, managing context, utilizing various tools, and dividing tasks among subagents. OpenAI's announcement documentation states that agents can work with files, run code, save temporary results, and operate in long-lasting sessions.
From "Answering" to "Completing"
Practically, agents change the unit of work from a single interaction to a series of tasks. In an internal report published by OpenAI, the company stated that the use of Codex is increasingly shifting towards longer-duration tasks. In May 2026, 70.2 percent of analyzed users submitted at least one task that was estimated to take more than one hour if done by a human. This figure comes from OpenAI's Codex usage data, so it should be read as a signal of adoption within their ecosystem, not a representation of all workers.
Examples can be found in several fields:
- Engineering: analyzing logs, searching for potential causes of errors, and preparing code changes for review.
- Operations: combining sales data, customer tickets, and inventory to identify issues that need prioritization.
- Finance: checking unusual transactions and creating a list of cases that need human review.
- Research: searching for documents, comparing findings, running experiments, and drafting initial reports.
OpenAI also reported that the use of agents is growing beyond technical teams. Departments such as legal, finance, and recruiting are said to be starting to use Codex as a primary tool for some of their work. This fact is interesting not because all jobs will soon be replaced, but because the boundary between "technical" and "non-technical" work is becoming more fluid.
The Issue Is Not Just Whether Agents Are Smart
When AI only generates text, errors usually appear as incorrect answers. When AI is given access to systems, mistakes can turn into actions: sending emails, altering data, opening tickets, running deployments, or sharing documents in the wrong places.
That’s why agents should be treated like fast junior colleagues, not like machines that are always right. They can perform many steps, but they may not understand business consequences, authority limits, or context that is not explicitly written in the system.
The greater the agent's ability to act, the more important it is for humans to define what the agent should not do.
Another risk is errors that seem reasonable. An agent can create analyses with a neat structure but use incomplete data. It can find patterns in dashboards but misunderstand metric definitions. Even when data access is restricted, results still need to be checked because access permissions do not automatically guarantee correct interpretation.
New Trend: Evaluation Happens While AI Works
Another development worth noting is the increasing focus on independent evaluation. On September 18, 2026, Anthropic announced a partnership with Accenture to conduct evaluations and red-teaming of AI models. The goal is not only to test model capabilities in the lab but also to see how systems behave when used in real work environments.
According to Anthropic, independent evaluators can help make safety claims easier to verify, although the responsibility remains with the model's producing company. This indicates an important shift: AI testing cannot be done merely through a few example questions. Systems need to be tested under conditions that resemble actual use, including data access, tools, conflicting instructions, and potential recovery when errors occur.
What Does This Mean for Us?
For everyday users, AI agents may not immediately take over daily tasks. However, their working patterns can already be applied to low-risk tasks, such as organizing documents, comparing several sources, drafting reports, or transforming raw data into action lists.
For teams and companies, the more important question is not "which agent is the smartest?" but:
- What data can the agent read?
- What actions should only be prepared but not executed automatically?
- When must humans approve the results?
- How can we track the sources of data, decisions, and changes made by the agent?
- What is the recovery procedure if the agent takes a wrong step?
OpenAI provides another example through the Data agent in ChatGPT Work, which can connect company data, investigate metric changes, create dashboards, and propose actions through connected tools. Systems like this demonstrate the real benefits of agents but also show why metric definitions, access rules, and trusted data sources must be established first.
What Can Be Done Now
- Start with cancellable tasks. Choose work such as drafting, grouping data, or compiling recommendations. Avoid giving direct access to payments, data deletion, and external communications.
- Separate reading and acting. Let the agent analyze data first. Actual actions should wait for human approval.
- Use limited environments. For technical experiments, use a sandbox or special accounts with minimum permissions. Do not grant administrator access directly.
- Keep a record of work. Document inputs, tools used, changes made, and reasons for approval. This log is important when results need to be audited.
- Test with failure cases. Provide incomplete data, conflicting instructions, and denied access conditions. A good agent is not just one that succeeds when everything goes smoothly, but one that safely stops when in doubt.
AI agents can indeed save time, especially for lengthy and repetitive tasks. However, the real value does not come from letting agents work unchecked. That value emerges when humans design clear boundaries, provide the right context, and place checks at the right points.
Thus, new skills in using AI are not just about writing long prompts. What is more important is designing work: defining goals, access, action limits, success evidence, and recovery paths. That is where the difference between helpful automation and automation that creates problems lies.
Sources & Further Reading
- Introducing the Agents API — OpenAI
- How agents are transforming work — OpenAI
- Partnering with Accenture on embedded evaluation — Anthropic
- Now everyone can put data to work — OpenAI
– Rio Yotto @rioyotto
