Fast Agentic AI Adoption Through a Structured Deployment Approach

At an event I attended a few weeks ago, I heard a respected industry panellist say that, on average, only 11% of GenAI projects worldwide make it into production. I found this statistic extraordinary, especially given how eager companies are to adopt agentic AI for business optimisation – with an apparent adoption rate of 60 – 70% – and the poor return on investment from a technology that dominates the news and has captured everyone’s imagination.

The pattern is remarkably consistent: a promising proof of concept, then a pilot, then another pilot, and somewhere along the way the business case quietly slips into next year.

In a recent webinar, my co-founder and CTO, Guy Ernest, and I explained why this happens and what we believe is a better way to build. This article is my written take on that discussion. The short version is that we start with the wrong part of the architecture – the AI agents – and create code as if there were no tomorrow (because much of it is now automated), leaving the most important architectural elements until last.

Why pilots never ‘take off’

Everyone is excited about AI agents, and there is a great deal of noise around them. Because everyone is talking about agents, teams start with the agents. AI-assisted coding has made development cheap, so they choose a framework such as LangChain, LangGraph, AutoGen or PydanticAI and produce vast amounts of code. Almost all the attention goes into designing and coding the agent.

In a recent discussion with a CTO whose organisation had already started down this path, I was told that his team had created an unmanageable and unverifiable codebase exceeding one million lines. That is easy to do when AI generates the code, but rather excessive when the goal is business automation through AI agents.

Beyond this excess, an agent is useless until it can access your company’s data and the systems your people use every day. That access is almost always an afterthought. Teams either borrow tools of questionable quality or hand-code their own, connecting them directly to the agent with no guardrails, security or observability. Tools are deterministic objects that interact through complex JSON data structures; LLMs are non-deterministic and can struggle with those structures, so connecting the two is always challenging. An LLM can call a tool incorrectly, fail to create or interpret the required JSON structure, or simply ‘hallucinate’. Such a hallucination could instruct the agent to do something damaging, such as deleting a database or important records, or emailing the wrong people.

The Model Context Protocol (MCP) was created to address these issues by wrapping tools in guardrails and providing the ‘intelligence’ required to make them suitable for AI agents and the non-deterministic nature of LLMs. However, MCP is often bolted on near the end of the architecture and design process, deployed poorly and rarely thought through. Teams also tend to adopt questionable third-party MCP servers rather than building properly with well-defined APIs and code mode.

The upshot is that the same issues keep arising:

  • Security, observability and scalability comes last. It is costly, so it gets built after everything else.
  • Tool misuse. Hardwired tools mean agents can misuse them, ignore them or hallucinate their way into real damage.
  • Humans left out. Approvals, notifications and visibility for people are rarely designed in.
  • Agentic complexity. Complex tasks require intricate execution graphs paired with equally complex prompts.
  • Evaluation becomes a last-minute panic. How do you prove that the system delivers the required business outcome consistently, day in and day out, like a good employee? Not by eyeballing it. Without the right platform, proper evaluation requires yet more coding.

The result is a never-ending cycle of rewriting, retesting and re-evaluating while security is retrofitted so that the CISO can finally say yes. LLM usage balloons, the return on investment recedes further into the future, and even when something ships, staff do not adopt it because it remains peripheral, isolated and difficult to control. The business case never materialises.

Build it like a house

We propose five steps, in this order:

  1. MCP servers – the foundations. Wrap the IT systems your people use in secure, scalable and observable MCP servers so that your technology estate is ready for AI.
  2. Agents – the walls. Layer agents on top without writing a single line of code. Done well, they operate in a tested, secure environment and require only business-language prompts.
  3. Agent teams – the rooms. Complex processes need several agents working together, not one heroic agent. Done well, these teams should be self-organising rather than dependent on complex, predefined execution paths and guidance.
  4. Channels – the roof. Connect everything securely to Teams, Slack, email, Discord or WhatsApp, where your people already work.
  5. Evaluation – the snag list. As with a finished house, you inspect the result and fix what needs attention using an objective, repeatable method.  Evaluation comes last chronologically, but it is never an afterthought: you plan for it from day one.

The point is simple: if the foundations are sound, everything above them becomes much easier – security, scalability, observability and manageability, alongside faster delivery.

Start with the foundations

This is the step most teams skip, yet it is the most critical. On our TrueX platform, each customer receives a private, single-tenant deployment located alongside their own data. This is not conventional SaaS, where something as sensitive as access to CRM or HR data resides in an environment the customer does not control.

In that environment, expect dozens of MCP servers, or connectors, because your organisation has dozens of data systems. Each server is a governed doorway with its own policies: which tables can be queried or updated, and what can be read, written or deleted. Document-based knowledge can be exposed in the same way through MCP RAG servers.

Agents, teams and channels

On top of this foundation, an agent is no longer copious amounts of code. It consists of three things: instructions, an LLM and the connectors it is permitted to use. The instructions are written in plain business language, so the task calls for a business analyst who understands its complexity, not a developer. Administrators approve the available LLMs—Claude, Google, DeepSeek, Kimi or whatever makes sense for cost and accuracy—and one is selected for each agent. Not every task needs the most expensive model, especially when the tools are good. Once the plumbing is solid, the main concern becomes controlling LLM spend, which is a far easier problem.

Simple tasks suit a single agent; difficult ones do not. A lone agent starts making mistakes and will never admit to them. Complex tasks need teams of AI agents that share data and communicate with one another. These teams should be self-organising and work towards completing the assigned task. Every call should be traced so that you can see where the team succeeds, where it slips and how to tune it until you can say, ‘I trust it.’ Producing a quotation, reviewing a proposal or analysing a customer request are more approachable examples of the same pattern.

Then comes the human dimension. AI agents assigned tasks that carry business risk should run on triggers and communicate through channels the organisation already approves. They should request human approval through those channels, much as quality assurance governs the work of employees today.

Prove it, do not eyeball it

Evaluation should be planned from day one and evolve alongside the rest of the AI solution, integrated at the data level. Evaluations should run frequently to validate accuracy against desired business objectives and outcomes, as well as cost, scalability and security. Data and agentic AI activity should be assessed not only by comparing outputs with ‘ground truth’, but also through modern LLM-as-a-judge methods.

Objective, defensible evaluation based on sound methodologies is essential from the outset and must continue beyond go-live.

Where should you start?

If you are stuck in pilot mode, my challenge is simple: stop building yet another agent. Choose the system your people use most, wrap it properly in secure, observable MCP access, and let them use it through the chat tools they already know. Everything else builds on that foundation.

This is the approach we have built into TrueX: helping organisations move from idea to production without becoming trapped in the pilot phase.

Scroll to Top