HomeAboutWhat We DoOur WorkInsightsVenturesContactStart a Conversation
The Mok Company
Back to Insights

The Mok Company

Where AI Saves Time, and Where It Quietly Costs You

What the evidence says about AI and time: real gains on defined tasks, slower work outside AI's frontier, and why individual gains rarely reach the P&L.

By Mohamed Mokhtar7 min

Most leadership teams now hear two stories about AI at work. In one, it saves hours every day. In the other, it is expensive noise. The research supports neither story on its own. It shows large gains on some tasks, worse results on others, and a gap between what individuals feel and what the business books. This article sets out what the studies found, how far each one can be trusted, and what that means for operations, customer experience and executive teams in Egypt and the wider region. We have kept to numbers we could trace to a source, and we state each study's limits, because those limits are the useful part.

The evidence for real gains

The strongest results come from work that is defined, repeated and written. A study of 5,179 customer-support agents at one firm found that an AI assistant raised productivity, measured as issues resolved per hour, by 14% on average [1]. The gain was 34% for novice and low-skilled workers, with minimal impact on the most experienced. One reading is that the tool narrows the gap between newer and experienced staff, which is useful for any team with high turnover or a long ramp-up.

Writing shows a similar pattern. In an experiment with 453 professionals doing writing tasks, average time fell by 40% and output quality rose by 18% [3]. In a preregistered experiment with 758 consultants at BCG, those working on tasks inside AI's capability completed 12.2% more tasks, and did so 25.1% faster [2]. A randomized trial with 95 developers found that those using GitHub Copilot completed a coding task 55.8% faster [4].

Read these figures with their limits in view. The support study covers one firm. The writing and consulting results come from controlled tasks, not from a normal quarter of messy work. The Copilot result comes from a single, well-defined coding task, and the sample was small. The models are also older than the ones teams use today, so the numbers describe a direction, not a forecast for your business. None of that makes the findings unusable, and a careful leader should not dismiss them. Still, the direction is consistent: where the task is clear and the output can be checked, AI can help, and it helps newer staff most.

The jagged frontier: where AI makes work worse

The same BCG experiment contains the finding that matters most for risk. On a task outside the frontier of what the AI could do well, consultants using AI were 19% less likely to produce correct solutions [2]. The researchers call this a jagged frontier: AI can handle one task superbly and fail at a similar-looking one, and nothing on the screen tells the user which is which.

A second result points the same way. METR ran a randomized trial with 16 experienced open-source developers working on their own mature codebases. With AI tools they took 19% longer. Yet they believed AI had sped them up, by about 20% [5]. The sample is small and the tools were early-2025 versions, so it would be wrong to read this as proof that AI slows experts down everywhere. What it shows is narrower and more useful: people are poor judges of their own time savings, and experienced staff on complex, familiar work may gain the least.

For a leader, the practical risk is a quiet one. Work that feels faster can be slower, and work that looks finished can be wrong. If your only evidence is how your team feels about the tool, you do not yet have evidence.

Customer experience teams should take this seriously. A fluent, confident reply that is wrong costs more than a slow, correct one, because the customer acts on it. The same applies to a summary an executive relies on or a figure that enters a report. The question to ask of every use case is not only how much faster it is, but what it costs when it is wrong and who would notice.

Two-column evidence card. Where AI helped: 14% more customer-support issues resolved per hour, 34% for newer agents, 40% less time on writing tasks with 18% higher quality, 25.1% faster inside AI's frontier, and 55.8% faster on one coding task. Where it hurt: 19% less likely to reach the correct answer outside the frontier, and 19% slower for experienced developers who believed they were about 20% faster. Each figure carries its source number.
Selected findings from sources [1] to [5]. The studies measure different things, so the figures are not directly comparable.

Why individual gains don't show up in the P&L

If the task-level gains were large and widely applicable, they should appear in wages, hours and output. One of the broadest studies looked for exactly that. Using Danish administrative data, Humlum and Vestergaard found precise null effects on earnings and recorded hours, ruling out effects larger than 2%, two years after ChatGPT's launch [6]. That is one country and one period, and it measures the whole labour market rather than a well-run workflow, so it does not say AI cannot pay off. It says that spreading tools across a workforce has not, by itself, moved the numbers.

Adoption is not the limiting factor. Stanford's 2026 AI Index reports that organizational adoption of AI reached 88%, and that documented AI incidents rose to 362 [7]. Wide use and a growing record of things going wrong sit side by side.

Our reading, and it is a reading rather than a finding, is that individual time savings leak away before they reach the business. Minutes saved on a draft are spent checking it. A faster first step meets the same approval queue. Gains go to the people who were already quick, or are absorbed by more output nobody asked for. Unless the workflow around the person changes, a saved hour is just a slightly different hour.

There is also a cost that rarely appears in a business case: the time spent reviewing, correcting and governing AI output. It is real work, it is often done by the most senior people, and it is the first thing a pilot leaves out when it reports time saved.

What actually works: workflows, measurement and human checkpoints

This is where The Mok Company's method sits, and we describe it as practice rather than as results. We do not quote client numbers here because the point is the approach.

First, start from a workflow, not a tool. Choose a repeated process with an owner, such as resolving a service request, preparing a proposal or routing an inventory decision, and map where time and rework actually occur. Test the AI at that step and nowhere else at first.

Second, measure before and after on the work itself. Agree the baseline in advance: time to resolution, rework rate, error rate, cost per case. Do not rely on self-reported savings. The METR result is the reason why.

Third, place human checkpoints where the frontier is uncertain. Decide what the system may recommend, what it may execute and what always returns to a person, in proportion to the cost of an error. The BCG result is the reason why. Begin with a narrow boundary people can observe, and widen it only where the evidence supports it.

Fourth, change the process around the tool. If a task gets faster, decide what happens to the freed time, and remove the approval step or handoff that would swallow it. This is what moves a result from a person's day to the operating numbers.

In practice this means a small number of well-chosen workflows, each with a named owner, a baseline, a decision boundary and a review rhythm, rather than a broad rollout that no one can evaluate. It is slower to announce and faster to learn from. If a workflow does not show a gain against its baseline, we would rather stop it than scale it.

The Egypt and MENA context

Almost all of the evidence above was produced outside Egypt and the region, and none of it was designed to measure organisations here. We would not assume the percentages transfer. Language mix, data quality, legacy systems and approval culture can all change the result, and in Arabic-language customer operations they are worth testing directly.

The policy direction is clear. Egypt's National Artificial Intelligence Strategy (2025-2030) sets targets of a 7.7% GDP contribution, 30,000 AI specialists and more than 250 AI startups [8]. These are targets, not outcomes, and they signal where talent, funding and attention are heading rather than what any one company will gain.

For mid-to-large companies in the region, this argues for a disciplined start rather than a late one. Pick one workflow where you can measure honestly, learn how AI behaves on your language, data and customers, and build the checkpoints before the volume arrives. Being early only helps if you can tell whether it is working.

Key takeaway

AI saves real time on defined, checkable tasks and can cost time or accuracy just outside them. The gains reach the P&L only when a workflow is redesigned around them, measured against a baseline and protected by human checkpoints.

Sources

  1. Generative AI at Work — Brynjolfsson, Li & Raymond, NBER Working Paper 31161, 2023
  2. Navigating the Jagged Technological Frontier — Dell'Acqua et al., Harvard Business School Working Paper / SSRN, 2023
  3. Experimental evidence on the productivity effects of generative AI — Noy & Zhang, Science, 2023
  4. The Impact of AI on Developer Productivity: Evidence from GitHub Copilot — Peng et al., arXiv, 2023
  5. Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity — METR, 2025
  6. Large Language Models, Small Labor Market Effects — Humlum & Vestergaard, NBER Working Paper 33777, 2025
  7. AI Index Report 2026 — Stanford HAI, 2026
  8. Egypt National Artificial Intelligence Strategy (2025-2030) — OECD.AI Policy Navigator

Test one AI workflow with evidence

See how The Mok Company approaches enterprise AI through measured workflows, clear decision boundaries and human accountability, or talk through a workflow you want to test.