Alibaba.com tests its AI agents on 107 commerce tasks. Getting the work done is only the first test.
The company also asks how long the work takes and what it costs. An agent can finish a task and still be too expensive for a customer to use every day.
For Alibaba.com, the global platform connecting business buyers and suppliers, that is a practical constraint. Accio Work, its agent workspace, helps small and midsize businesses (SMBs) with sourcing, supplier communication, and ecommerce operations. Those customers need the work done at a price they can sustain.
Kuo Zhang, president of Alibaba.com, puts the constraint directly:
“Only affordable AI is a useful AI for SMBs.”
In Episode 244 of The Artificial Intelligence Show, part of our AI Transformations series, Zhang explains how Alibaba.com tests agents on commerce work and compares completion, time, and price.
Broad measures of model ability do not answer every question a business has about an agent.
Zhang describes the gap between strong performance on a subject such as mathematics and success in a messy commercial situation. The latter involves particular requirements, real business information, and tools that the system must use to get something done.
So Alibaba.com built Commerce Agent Bench to test that kind of work. Zhang says the company drew on millions of conversations between buyers and sellers, along with business workflows, to create 107 benchmark tasks. The set is intended to grow as the company adds more complex business problems.
Zhang describes task completion as a binary test: An individual task either passes or fails. Across a set of tasks, the pass rate shows how often the agent succeeds at doing these tasks.
Leaders can learn from this, even if they don’t have an elaborate benchmark for their agents. When evaluating agents and AI, choose recurring work that matters, specify the required result, and decide what would make that result acceptable. If the team cannot agree on what completion means, it will struggle to decide whether an agent is worth scaling.
A supplier comparison illustrates the point. A useful test would ask whether the result addresses the buyer's actual requirements and contains the information needed for a decision. A neatly formatted table by itself cannot answer that question.
Completing the task establishes that the agent can do the work. It does not settle whether that system is the right choice for repeated use.
Alibaba.com also considers how much time and how many steps an agent takes to reach the result, along with the price of doing so.
And the stakes change as use expands. A cost that looks acceptable in a trial may look different when a small business relies on the system daily, or when a larger company runs it throughout a department.
Completion also changes how cost should be interpreted. Suppose an agent costs less to run but needs repeated attempts or staff cleanup before its output is usable. Its apparent price advantage may disappear. A more expensive system could be the better choice for that job.
That suggests a more useful management question:
What does it cost, in money and time, to get an acceptable result at the volume we need?
This is a question for the people who own the workflow as well as the people choosing the technology. The business owner can identify the required result, the tolerable delay, and the cost of getting it wrong. Those constraints help determine which system is suitable.
Once evaluation starts with the work, a single company-wide model choice becomes less obviously sufficient.
Zhang describes testing different models inside the same execution system used by Alibaba.com's agents. Keeping that surrounding system consistent lets the company compare how the models handle the same business problems.
“Routing to the best model fit for the right problem [is the goal,” he says.
The practical implication is that the appropriate choice may differ across tasks. A business can examine whether a particular job benefits enough from a more capable or more expensive model to justify using it there.
Alibaba.com also compares complete agent systems on the same tasks. That is a separate evaluation because an agent's performance depends on more than the model. The surrounding software determines which tools it can use and how it carries out the work. The available business context shapes whether it has the information needed to succeed.
For leaders, that distinction helps locate the next investment. A weak result could reflect a model limitation, missing information, or an inability to take the required action. Buying more model capability will help only when it addresses the constraint.
At the end of the day, a company's own test set is most useful as a way to make decisions about its work. Alibaba.com's commerce tests reflect its customers, systems, and business context. Another organization needs to identify the tasks and conditions that matter in its own operations.
Those conditions will also change. Zhang says Alibaba.com expects to keep expanding its task set. As agents take on more complex work, the tests need to reflect what they are being asked to do.
Before approving wider use, leaders can answer four questions:
Before approving wider use, ask the workflow owner what it takes to get an acceptable result, including the human work needed to finish it. Then test that cost against how often the business needs the result. That can justify expanding the deployment, changing models, or fixing the workflow before buying more capacity.
A model's price tells you what it costs to run. The scaling decision depends on what it costs to finish.
Did you enjoy this transformation story? Go deeper on how AI is reshaping work and business with The Artificial Intelligence Show. Each week, we break down what matters in AI and what leaders should do about it.