Insights/Product

More Than an LLM

·Davide Martucci

I joined Lawrence Baker on the Investment Association's IA Talks AI podcast to talk about deploying AI inside one of the most regulated industries in the world.

Listen to the episode: Building Trusted AI: Why Investment Operations Need More Than an LLM

Below is the longer version of what I said.


Automation amplifies whatever it sits on

A senior executive running trading operations at a tier-one bank put the problem to me plainly. He said the presentations were remarkable, the pain points were real, and yet it all felt like a toy he could not play with. He did not know how to contextualise it in his industry.

He was right, and the reason is structural. Two barriers have defeated every generation of automation technology in our industry, and neither of them falls because a new model was released.

The first is data. Information arrives from custodians, fund administrators, counterparties, market data providers and internal systems, in different formats, on different schedules, with different levels of completeness. The second is process complexity. Every firm has its own thresholds, its own escalation logic, its own regulatory interpretation, its own client-specific arrangements. These are not a standardised assembly line. They are a web of hyper-customised workflows built over years, across jurisdictions, by many hands.

Automation amplifies whatever foundation it sits on. Build on chaos, and you get automated chaos. A capable agent sitting on fragmented data, with no deterministic execution layer, no human handoff and no audit trail, is not a solution. It is a liability with better prose.

Generic AI is abundant. Contextualised AI is rare. We set out the full version of this argument in Why Agentic AI Needs the Right Ecosystem.


Seven years before the agents arrived

Next Gate Tech has spent seven years building enterprise SaaS for investment operations. Reconciliation, NAV oversight, fees analytics, compliance monitoring, referential data, entity management. Production workloads for BNP Paribas Securities Services, Amundi, Eurizon and others, across more than €500bn in assets.

When agentic AI arrived, we did not start again. We transformed what we already had into an ecosystem for agentic workflow automation.

Our analytics became NGT Tools: the same validated computations our clients have relied on for years, now exposed as tools an agent can actually call, through MCP and APIs, scoped to what its role permits, with every invocation attributed and logged.

Our operational knowledge became NGT Skills: how NAV oversight fails in production, how fee methodologies vary across jurisdictions, what post-trade reconciliation really looks like at month-end, how escalation has to be structured to satisfy a CSSF auditor versus one from the Central Bank of Ireland. Packaged so a model can apply that knowledge rather than approximate it.

And the workflows themselves. Seven years running live operational workflows for clients across the full value chain of investment operations, at volume, across large numbers of funds, vehicles and entities. Reconciliations that clear every morning. NAV validations that run before release. Compliance checks that have to return the same answer twice. Today more than 10 million entities move through those workflows daily.

That orchestration layer is now part of the ecosystem. Agents can invoke those workflows, and agents can build new ones. Orchestration is where the automation actually happens, and it is the part that took seven years to earn.

And beneath all of it, the data foundation. Harmonised from every source across the value chain, not just the internal systems of one firm but the full network of counterparties, service providers and data sources that operations depend on. Business identities resolved. Relationships mapped. Continuously updated. Single tenant and fully sovereign, never crossing into shared infrastructure or external model training.

That is what grounding a model means. Not a better prompt. Certified data beneath it, real tools in its hands, domain knowledge it did not have to guess at, and a workflow engine that turns its decisions into work that completes.


Rules set once, inherited by every agent

Grounding makes one agent capable. Controlling a hundred of them is a different problem.

The moment agents become easy to create is the moment they become easy to lose track of. Different teams, different credentials, different escalation logic, none of it visible from a single place. Shadow agents are the new shadow IT, and in a regulated environment that is not a productivity story. It is an audit finding waiting to happen. We called this agent sprawl in Build With Agents. Run Without Them, and it is the most common pattern we meet: pilots multiply quickly, consistent control does not.

So SPARK puts a control layer above the agents. One workspace where the rules are set once. Which data an agent can see. Which NGT Tools and NGT Skills it can reach. Which actions wait for a human signature. Who owns it. What evidence is retained. Credentials issued from a central vault, scoped per role, rotated centrally, never living in a prompt or a script. Revoke the agent and you revoke its access everywhere.

Defined once, inherited by every agent you ever run. Including the ones you will build next year, on models that do not exist yet. Beyond the Agent Harness explains why that layer has to stay yours rather than living inside any single model or vendor.

We do not govern agents to constrain them. We govern them so we can deploy more of them with confidence.


The operation runs on triggers

Operations teams ask a simpler question than any of this: what happens at four in the morning when nobody is watching.

A custodian file lands. A NAV moves past threshold. A break ages past SLA. A limit breaches. A GP document arrives.

The right agent wakes, reasons over certified data, calls the NGT Tools its role permits, opens a workflow, assigns an owner, and stops where a human decision belongs.

And where the work repeats, the agent does not need to be there at all. Agents are extraordinary at design time: creating the workflow, mapping the exceptions, translating a messy operational requirement into a precise process. Then the workflow runs deterministically, with no model call, reproducible and auditable and cheap.

The agent builds the engine. The engine runs without the agent.


Human on the loop

The most effective deployments we see are not fully autonomous, and were never meant to be.

Agents handle the volume and the pattern work. They surface the right information to the right person at the right moment, with enough context to act immediately. The operations professional sees exactly what the agent did, why it escalated, and what it expects them to decide.

Agents own progress. People own decisions.

Human on the loop is not a safety feature bolted on at the end. It is a design principle, and the handoff deserves as much engineering as the automation. The people spending their days on exception queues were not hired to do that. They were hired to think, to judge, to oversee. We are not building AI to replace judgment. We are building AI to restore it.


We measure what the agents do

Trust has to be demonstrable, so we measure it.

Agents in SPARK are versioned, test-run and measured before and after every change, using tasks, trials, traces and evaluators that make behaviour visible from design through to production. We wrote about that discipline in Building Agents We Can Measure.

A model swap, a prompt change, a new tool: in most organisations these ship on intuition. In a regulated operation they need evidence, before deployment and after. Otherwise you are not running agents. You are hoping.


What you get is an output you can trust

All of this only counts at the point where something comes out of the system.

No board asks how many agents you are running. They ask whether the NAV was right, whether the breach was caught, whether the report that went to the regulator holds up. The only thing that counts is what comes out of the system and whether you can put your name on it.

That is what SPARK is built to produce. A NAV validated against certified data, with every check reproducible. A break investigated, root cause identified, owner assigned, cleared. A compliance result with the evidence attached rather than promised. A fund record you can certify because you can see how every field got there.

And behind each one, an answer to the question a regulator will eventually ask: what was this agent allowed to do, who decided that, and can you show me the run.

Every run logged. Every run traceable. Every run replayable.

The goal was never maximum autonomy. It is the right autonomy.


The right autonomy

Moving beyond experimentation is not a matter of finding a better model. The models are already good enough. What is missing in most institutions is everything around them: the data they stand on, the tools they can safely call, the knowledge they should not have to invent, the evidence that they work, and the control layer that makes the whole estate accountable.

Investment operations need more than an LLM. They need an ecosystem that grounds it and a control layer that governs it, so that what comes out the other side is an output you can trust.

Listen to the full conversation on IA Talks AI


Next Gate Tech builds SPARK, the operating ecosystem for agentic AI automation of investment operations. Clean data, deterministic workflows, domain knowledge, and full traceability. Purpose-built for one of the most heavily regulated industries in the world. Trusted by BNP Paribas Securities Services, Amundi, Eurizon and others across €500bn+ in assets.

nextgatetech.com