What happened

At Beijing’s World Humanoid Robot Games on 22 August 2026, a robot completed a 100-metre sprint in 9.39 seconds, according to Associated Press reporting. That time was below Usain Bolt’s human record of 9.58 seconds.

It is an eye-catching demonstration. It is not, on its own, a measure of readiness for a workplace.

Why it matters

A demonstration is usually built around a clear task and a visible result. Business operations contain interruptions, incomplete information and competing priorities. The difference deserves attention whether the system moves through a warehouse or works inside a spreadsheet.

Suppose a company is assessing a robot for an internal delivery route. Travelling quickly is useful only if the machine reliably collects the correct item, reaches the correct destination, handles obstructions and stops safely when conditions change.

The same logic applies to a software agent completing a purchase request. Creating an order quickly does not show that it selected an approved supplier, recognised a duplicate or respected the available budget.

I would make the acceptance test cover the complete job. Include the preparation and recovery work performed by people. Otherwise, the impressive part of the demonstration can hide the parts that determine whether the operation is worth deploying.

The bigger shift

A useful evaluation programme needs ordinary cases and awkward cases. Define both before watching the product demonstration, so the test does not simply reward whatever the system happens to do well.

For a physical system, the trial might examine a blocked route, an unexpected object and an operator requesting an immediate stop. Appropriate safety specialists should define the environment and conditions. A controlled trial should not expose bystanders to an experiment.

For a digital workflow, the equivalents could be conflicting records, an unavailable application or a request outside the agent’s authority. Test what the system does when it cannot finish. A reliable refusal or escalation can be a successful result.

Record interventions separately from completed tasks. If a person repeatedly rescues the workflow, that time belongs in the performance assessment. Record recovery time too: how long until normal service resumes after something goes wrong?

Link the measures to the business case. For guidance on that starting point, see why AI strategy beats AI tools.

My take

A stopwatch gives a clear answer to a narrow question. Executives should be equally precise about the question their own evaluation is answering.

Before approving a pilot, write down what would make the system useful in the intended setting. Include the acceptable error rate, the required human support and the conditions under which it must stop. Decide who can expand or suspend the trial.

Then watch a full work cycle. Follow an item, request or decision from its arrival to its accepted outcome. Ask staff where they intervened and what knowledge made that intervention possible.

The demonstration can earn a place in the evaluation. It should not substitute for it.

For a business, readiness means that the surrounding process can absorb the technology’s limitations while delivering a worthwhile result. Speed is one contribution to that judgement, alongside reliability, recoverability and the people needed to keep the operation working.