Tool selection · May 2026 ·

Choosing AI tools by value per unit of effort

Vendor decks compare features. A buyer needs a decision. For a firm weighing Copilot against ChatGPT, we scored ten capabilities on value and on deployment effort. The winner was different for nearly every row.

The client had rolled out Microsoft Copilot across the organisation as the first step of its AI programme. Within months the next question arrived, as it tends to. Invest further in Copilot Studio, add ChatGPT Enterprise, or wait for frontier tooling to become enterprise-ready in the EU?

Every vendor deck answered with a feature list. The client needed a decision, though, and a feature list does not contain one. What a decision needs is a way to weigh what a capability is worth against what it would actually cost to run.

Method

So we tested both platforms against the workflows the firm actually runs: reading tender packs, matching people to requirements, filling templates, routing approvals. Every capability got two scores. Value, from one to five, for what the capability is worth if it works. And effort, again from one to five, for what it takes to deploy and operate. The index is simply value divided by effort. That is the whole method.

The simplicity is deliberate. The point was never statistical sophistication. The point is that every claim in a vendor comparison gets forced into two auditable judgments, which a reader can then disagree with one at a time.

ChatGPT Copilot 1.0 0234 index = value / effort Read & parse documents Search enterprise content Reason over unstructured data Drafting & writing in tone File manipulation (xlsx/docx) Verification & gap-checking Human-in-the-loop approvals Event-driven triggers Audit trail & logging Overall agentic capability
Value-per-effort index by capability, from testing in Q2 2026. Above 1.0 a capability pays for itself. Where the dots overlap, the tools tied. Scores are judgments rather than measurements, and they decay as the market moves.

ChatGPT led wherever the work was messy. Reasoning over unstructured documents scored 4.0 against 0.67, for instance, and an agent that reads an RFP pack, matches people to requirements, and fills templates end to end could be built in days rather than weeks. Copilot, on the other hand, led on the plumbing: event-driven triggers, approval flows, native SharePoint grounding, and EU enterprise compliance today. We also tested frontier tooling separately. It was stronger still on raw capability. But several of its differentiating features were not yet enterprise-ready in the EU, and capability you cannot deploy is worth exactly nothing this quarter.

Conclusions

  1. Make ChatGPT Enterprise the primary agentic surface, with one named owner per agent.
  2. Keep Copilot active for light in-flow tasks inside M365. Stop further Copilot Studio investment, where the same outcome took weeks instead of days.
  3. Re-test the top workflows quarterly against the current market.

That last point is the one we would defend hardest. The scores in the figure were true in Q2 2026. Some of them will be wrong two quarters later, and that is fine. A recommendation that carries its own expiry date is more honest than one that pretends to be permanent.