// THE GROWTH CHAIR

Don't Build the LLM. Integrate the Components.

The architecture that survives audit.

Illustration for the essay: Don't Build the LLM. Integrate the Components.

The auditor asks what your model did at 14:32 on a Tuesday in March, on a mortgage application that was declined and is now sitting in front of the Financial Ombudsman. You either pull up the run, show her which component handled what, and walk her through the decision. Or you don’t have a system. You have a demo.

That moment, repeated across every regulated industry I have worked in, separates AI products from AI press releases.

Most boards still ask which large language model to bet on. Wrong question. One big model is a single point of failure. Opaque. Expensive. Brittle on the narrow tasks that actually matter inside a business. The accuracy headlines are written against general benchmarks. Your benchmarks are not general. Your benchmarks are: did the system route this clinical pathway correctly, did it triage this support ticket, did it size this account, did it flag this contract clause. On those, a small domain-tuned model usually beats the headline LLM. A chain of small specialised components beats either of them.

The expensive version of this mistake is taking one frontier model and training it on your entire business. Every contract, every ticket, every Slack channel, every customer record, fed into a single fine-tune. The pitch sounds great. The bill is huge, the system cannot be audited the way any regulator wants, and you are now locked into one vendor’s roadmap, pricing, and licence terms. When any of those change, and they will, the cost of unwinding is a board-level event.

AT&T have just put numbers behind this in public. They rebuilt “Ask AT&T” away from heavy general-purpose models and toward a mix of smaller, domain-specific ones for the bulk of their twenty-seven billion daily tokens. The reported saving is around ninety percent. They run agents that plan and execute narrow tasks. They back open-source telco-specific models because generic LLMs were not built for telecom data. They keep humans in the loop on anything that touches a customer or the network. Hybrid architecture. Multiple models. Best fit per task. That is not an AI strategy lifted from a vendor deck. It is what an operator does when the bill arrives.

The same shape shows up wherever the work is regulated and specialised. When I was at Sapio Sciences, SigmaticOS was built as an orchestration platform, not a single model. Over a hundred agents, each tuned to a part of the drug discovery loop, sat under an orchestrator. The orchestrator picked the right agent, handed it the right context, then folded the result into the next step. In silico work on one side, lab automation on the other, model training in between. The LLM was not the product. The orchestration was the product.

Lumeon proved the same point without using AI at all. A care pathway is not one decision. It is fifty, taken across clinicians, systems and patients over weeks. The platform routed each step to the right component: the rules engine, the messaging layer, the EHR integration, the patient-facing app. Deterministic by design, because clinical care has to be. Every step inspectable. Every decision traceable to its source.

That is the bar AI deployments in regulated domains have to clear. Probabilistic components can sit inside this kind of architecture. They cannot become the architecture.

Here is the part most teams still miss.

The architecture above is necessary. It is not enough. What makes it actually work inside a regulated business, what makes it survive a board paper, a SOC 2 audit, an FCA inspection, a customer escalation, is the audit trail. Every call to every component, every input, every output, the version of the model, the version of the prompt, the data the agent saw. Replayable. Searchable. Owned by you, not the vendor.

Suresh Rajashekaraiah and others have started naming this layer as a category in its own right: observability for agentic AI. Agentic systems do not fail the way traditional software fails. The failures are not stack traces. They are emergent and probabilistic.

You cannot debug what you cannot see. You cannot govern what you cannot replay.

Boards that have lived through a serious incident understand this in their bones. Boards that have not are still treating observability as a cost line.

The pattern is not complicated. It is just unfashionable. Stop asking which frontier model to license. Ask three questions instead.

Which tasks are narrow enough that a smaller, domain-tuned component would beat a general LLM. The honest answer is: most of them.

What does the orchestration layer look like, and who owns it. If the answer is “the model vendor,” your moat is rented.

What is the audit trail, where does it live, and can a non-engineer reconstruct any single decision the system made in the last twelve months. If the answer is no, you do not have a regulated AI product. You have a liability waiting for a Tuesday in March.

The companies that win the next phase of enterprise AI will not be the ones with the biggest model. They will be the ones who can show their working.

// Originally published on The Growth Chair · 29 Apr 2026 · Join the discussion on Substack

// THE GROWTH CHAIR BY EMAIL

One essay like this every week. Free. No paywall.

Delivered via Substack. Unsubscribe any time. See our privacy notice.

// GET IN TOUCH

Clarity when it counts.

If you have a board seat, a fractional mandate or a commercial reset coming up in the next 90 days, email directly. We'll book 30 minutes to see whether Ortent is the right fit.

Book a 30-minute intro call