It takes a few hours to build an AI agent. It takes six to seven months to get it into production. Manoj Saxena - the executive who commercialized IBM Watson - built TrustWise around one thesis: intelligence without control is not deployable. In this episode, he joins Craig Smith to explain why 95% of enterprise AI agent projects are stalling between pilot and production, and why the answer has nothing to do with the quality of the underlying models. The bottleneck, Saxena argues, is the absence of an entirely new class of infrastructure, something that can evaluate every tool call, every action, every output of every agent at runtime, in milliseconds, against the full stack of alignment requirements that govern what an AI is actually allowed to do inside a real enterprise.
TrustWise's AI Control Tower does that across all vendors and agent frameworks simultaneously, operating in live, sidecar, batch, or simulation mode and aligning agent behavior against six layers of requirements - from UN Human Rights frameworks down to individual customer SLA commitments - in 10 to 300 milliseconds per decision. The conversation covers demonstrated results (83% cost reduction, 40% safety improvement), the token consumption paradox that's making agentic AI far more expensive than expected even as token costs fall, and a milestone Saxena compares to the moment data traffic surpassed voice on AT&T's network: last month, for the first time ever, agent traffic on the internet exceeded human traffic. The episode closes with a preview of Genesis agents, TrustWise's next product, designed not just to prevent bad outcomes but to surface beneficial hypotheses by looking 95 moves deep into enterprise data, in domains like fraud detection and revenue leakage, in the way Deep Blue looked 95 moves deep in chess.
Subscribe to Eye on A.I. for weekly conversations with the people building and deploying the future of AI.
[00:00:00] [SPEAKER_01] Trust has become an issue. It's not simply trusting that the model is accurate, but trusting that the model is going to do what you want it to do and not what it wants to do.
[00:00:11] [SPEAKER_00] Last month, the first time ever, the traffic on the internet, agent traffic exceeded human traffic. Just like when on the AT&T network, data traffic exceeded voice traffic. This is a very big deal. What I saw was this focus on building more and more intelligent models and bigger models. Almost like you're building bigger and bigger nuclear cores, but no one's thinking about putting a dome on top of these things. Dashboard got me going with TrustWise about three, a little over three years ago.
[00:00:36] [SPEAKER_01] That's an interesting analogy, a dome over nuclear core. Why don't I have you introduce yourself to listeners and give as much of your background as relevant, of course, the IBM Watson thing and how you got to TrustWise.
[00:01:00] [SPEAKER_01] And then we'll talk about the runtime governance and the control layer over agentic AI systems and all of that.
[00:01:11] [SPEAKER_00] I'm Anuj Saxena. I'm the CEO and founder of TrustWise. This is my fifth startup. So some people call me a certified masochist. But the reality is that I love building things and love building companies and assembling people together to go after important problems. And I've been working on the intersection of enterprise AI, responsible AI and operationalizing trust in digital systems for a long time.
[00:01:41] [SPEAKER_00] Going back to the IBM Watson era about 10 years ago, I had the privilege of building a startup that was acquired by IBM. And then after that, the IBM board asked me to commercialize the Jeopardy playing game called Watson, which I yeah. So I had I spent about three years taking that game that was a size of a master bedroom.
[00:02:05] [SPEAKER_00] And I live in Texas is a big master bedroom and reducing it down to one pizza box level system that we deployed. And so, yeah, that's kind of been my journey. I was teaching responsible AI at the University of Texas, Austin, and seminars at Cambridge University in London when ChatGPT got launched. And what I realized, I was surprised how fast this technology had moved.
[00:02:31] [SPEAKER_00] And what I saw was this focus on building more and more intelligent models and bigger models, almost like you're building bigger and bigger nuclear cores. But no one's thinking about putting a dome on top of these things. I decided to put on a jersey and come back and said, this is a problem that needs to be solved. Otherwise, my grandkids are going to ask me saying you were there. You know, you made all this money with your companies.
[00:02:58] [SPEAKER_00] Why didn't you solve it? So that's what got me going with Trustwise about three, a little over three years ago.
[00:03:03] [SPEAKER_01] Yeah. And that's an interesting, interesting analogy, a dome over a nuclear core. You're focused on building a platform that sits above everything that is a control layer, both for generative AI and agentic AI. Is that right?
[00:03:28] [SPEAKER_01] And before we get into that, can you talk about why trust has become such a critical issue, trust and governance? And I'll just say that personally, you know, this is moving so fast, you know, now with Mythos, you know, from Anthropic, which is so powerful, they can't even release it.
[00:03:57] [SPEAKER_01] But the issue of trust goes in many directions. It's not simply knowing, trusting that the model is accurate, but trusting that the model is going to do what you want it to do and not what it wants to do. So can we talk about how trust has become an issue, how it's changed in the last year or so?
[00:04:26] [SPEAKER_00] Absolutely. I think you nailed it. At the end of the day, trust cuts down to does the AI and the model do what you intended to do? That kind of the heart of is it aligned to your business and personal intent? Does it, you know, so the alignment problem is not spoken about enough. And the reason trust has become more and more important is three things. Number one, AI has now moved from generating output to now taking actions, right?
[00:04:54] [SPEAKER_00] Watson and deep learning and chat GPT is not just giving you a better email or a better picture. Now it's taking actions on your behalf. And these actions could run for minutes and hours and days. So essentially, you've got digital workforce that's being introduced that can act now. Second, enterprises and enterprise AI is moving from a single model workflow to a multi-agent system. So the rise of things like OpenClaw.
[00:05:22] [SPEAKER_00] I think OpenClaw is going to be a thousand times more impactful than chat GPT was in terms of its impact on business and society. So businesses are now moving towards deploying not just single model, but multi-agent system, multi-model workflows that if things break, you could have real chaos on your hand. And third is policy in terms of driving the behavior of these systems is no longer something you can enforce only a deployment. It has to be continuously enforced at runtime.
[00:05:52] [SPEAKER_00] So it's not enough to say, well, you go do this. You've got to make sure that you're checking every tool call, every action, every output. It is staying on policy. So as a result, the architecture has now become the risk surface. You know, the whole architecture of how you're building an AI stack is become a risk surface. And that's why trust becomes like the silver thread that runs through your entire stack.
[00:06:17] [SPEAKER_00] And there needs to be a new class of infrastructure to manage and deploy these agents at scale. I sort of talk about it almost like an HR department of agents. You know, you would not deploy a company today with humans in it without an HR and finance department. And now you're about to introduce a whole bunch of digital labor with no HR and finance, with no drug testing, with no employee manuals on how to behave, with no performance appraisals.
[00:06:46] [SPEAKER_00] So we look at what we're doing at Trustwise. We call it the AI control tower. It's almost like the HR and finance department for hundreds of thousands of agents from multiple vendors.
[00:06:56] [SPEAKER_01] That layer of the control tower, is that in itself an agentic system?
[00:07:04] [SPEAKER_00] That's a great question. And yes, at the end of the day, when you have to scale the system, you'll need AI to control AI. But the difference is that the agentic system that we deploy, we call these guardian agents. But these guardian agents are built with human in the loop. And these are built with deterministic outcomes. So it's not probabilistic systems.
[00:07:27] [SPEAKER_00] So what we are looking at is a whole new class of agents and humans working together to be able to define, assess, control, and then optimize these systems as they're running. And they're not enough humans. And one of the large global companies told me by the end of this year, they'll have 100,000 agents in the company. And you're not going to have enough employees being able to onboard each of these manually and test it manually.
[00:07:54] [SPEAKER_00] So in the control tower, these guardian agents, there are two types of agents, guardian agents and Genesis agents. Guardian agents make sure that these agents are onboarded well and they're working well. And Genesis agents convert some of these agents into super workers, what is called as AGI. So within our control tower, there are two types of agents, guardian agents and Genesis agents.
[00:08:16] [SPEAKER_00] But most of the focus right now is guardian agents to make sure that the agents are aligned and are behaving properly so that you build the confidence to be able to then launch them.
[00:08:26] [SPEAKER_01] This is to provide runtime governance. I mean, it's monitoring what agents do in real time. Is that right?
[00:08:35] [SPEAKER_00] Yeah. So one of the issues is this is a wide open space today because runtime control means the control decision happens at the moment of action, not weeks before in a policy document or not hours later in an audit lot. So today, if you look at the problem, there is a wide space here. There are three types of software that enterprises have. None of them are able to manage runtime control.
[00:09:01] [SPEAKER_00] One is security and security mostly is outside in defense focus. Right. And agents are the new threat factors. Agents are the new insider threat. And there is no software for that. I like to say that I can build you the world's most secure prison and defend it from outside. But if you have a bunch of Chuckies and Hannibal Lecters on the inside, you're still going to have chaos on your site. So security doesn't protect it. Security is mostly outside in and not inside out.
[00:09:29] [SPEAKER_00] Second, governance only defines policies, but it doesn't implement policies at runtime. What I call as declarative governance versus runtime governance. And third, observability tells you that's the third class of software that companies have. It tells you what's going on, but it doesn't let you influence and shape it.
[00:09:48] [SPEAKER_00] So one way to think about it is you're launching the supercars into your company with a giant, you know, a thousand horsepower engine without any steering wheel, without brakes, without seatbelts and without emission control. So this is the area, this new layer. To give you an example, each of these agents that you deploy have to perform differently.
[00:10:13] [SPEAKER_00] So in the UK, there is a law called the UK FCA consumer duty law for distressed customer. So if you have a financially distressed customer, say an 85 year old who lost his wife is applying for a loan versus a 24 year old who just got a new job and she's applied for a loan. By law, the tone and clarity and helpfulness of your agent has to be different.
[00:10:37] [SPEAKER_00] Today, there are no systems that are able to differentiate and drive like all in the steering wheel to align the behavior of the agent at runtime to meet the requirements of those different individuals. So that's an example of runtime control, which means the system checks the action before it happens. Is this customer, you know, financially distressed? Is this action allowed? Is this tone allowed? Is there a human approval required?
[00:11:03] [SPEAKER_00] Is this action is not stopped or if it's something I don't have an answer to, do I say no to or do I escalate it to a human being? All of these things is what the runtime control implements at the time of action.
[00:11:17] [SPEAKER_01] Yeah. Yeah. And you have something called modular AI shields to watch agent actions and constrain tool use if necessary and ensure compliance with rules and regulations, as you mentioned. How do you differentiate this modular AI shield from a guardian agent? Are they the same thing?
[00:11:46] [SPEAKER_00] Yeah. Great question, Craig. So think of guardian agent as the AI that is, you know, overseeing and supervising different worker agents or functional agents. And that guardian agent can be equipped with one or more shields. And so it's like adding on a superpower pack. If you're Mario Bros., you're adding a superpower pack. And broadly, there are three types of shields. There are shields for safety and security. There are shields for cost and for compliance and regulations.
[00:12:16] [SPEAKER_00] And then there are shields for cost and carbon management. So typically, when you are guarding these worker agents, you are looking at combining all these three together. You want it to be safe and secure. You want it to be compliant with internal policies and external regulations. And you want it to be cost and carbon efficient. All three together is what we define as trust posture management.
[00:12:39] [SPEAKER_00] And that's the new space we believe is required is you need to be able to manage trustworthiness of these agents as they are acting at scale and be able to make those tradeoffs across safety, compliance and efficiency. And that's what the shields allow you to do.
[00:12:55] [SPEAKER_01] I see. And the shields are, they're not agents themselves. They're just watching. Who are they talking to when it is something that's questionable? Does it send it to the guardian agent? Does it send it to them?
[00:13:13] [SPEAKER_00] Yeah. The shields are always attached to a guardian agent. I see. So shields by themselves don't do anything. The shields get plugged in like a breastplate, like an armor. So a guardian agent will have multiple, you know, shields, which will give it the mission as to what it is allowed to do in terms of controlling your AI workforce.
[00:13:34] [SPEAKER_01] And this is all within a platform. How is, does the human interact? Because you said, depending on the policy or action, you may want a human in the loop.
[00:13:47] [SPEAKER_01] Is this a dashboard that, you know, sends alarms when a human needs to, or is it something more hands-on that the human is watching the agents perform?
[00:14:09] [SPEAKER_00] No, that's a great question. So the product is a platform, but the platform is a collection of APIs and what we call a CLI command line interfaces. So it is not yet another dashboard in a UX. It's a collection of APIs that you can call and orchestrate because people are not looking for one more pane of glass. That's right. I already have my security infrastructure. I want to plug this into security, into governance, into observability.
[00:14:37] [SPEAKER_00] So at the core level, the product is a collection of APIs and CLIs. But we do have a reference implementation of a dashboard. We call that the UX of the product. And that's where the humans will go in. And they primarily do four things. Number one, they onboard their agents and the workforce. So this is where you connect, you point to it, you discover it, you bring them on board.
[00:15:02] [SPEAKER_00] Number two, you then assess them and to say that you evaluate them. You say, okay, I'm going to put you to this job at a customer support board versus the other one. I'm going to put you to work for complaints handling. And the third one, you're going to do HR questions and answers. So based on that, we will then apply the right policies and controls. And we would run evaluations on it like a simulator, like a Formula One simulator in a car. You would run the agent through hundreds of thousands of flight paths.
[00:15:31] [SPEAKER_00] And then you will see where is it breaking, what policies and controls need to be changed. So one, you onboard, second, you assess and evaluate. Third thing you do is you then put that into production and you start monitoring for drift. Is it behaving right? Is it going off policy? What we call as rogue agents, rogue AI.
[00:15:51] [SPEAKER_00] Because some of these agents could go off policy and start doing things like they could get stuck in an invisible loop and eat up a lot of your GPUs and your compute tokens. Or they could start colluding against each other and come up with a language that is not English. It has even demonstrated that agents have figured out that when they talk to each other, they figured out that English is not the most efficient language. They discovered their own language. But if you do that, then you lose all the evidence and traceability.
[00:16:21] [SPEAKER_00] So we would then look for drift in those kinds of things. Is it drifting the behavior from what it's supposed to do? And then the last and final, it creates a layer of provable evidence for audit and learning to be able to say, okay, can you then now take these systems and prove, say, four years from now, if there is a lawsuit where your agent declined a claim and the patient died and you're getting sued.
[00:16:46] [SPEAKER_00] Can you go back and say on June 10th, you know, 2026 at 4 or 5 p.m., these collection of agents looked at this data and this particular claim and they decided this. They used these policies and this data set. And they came to this decision, which is now being sued. Can you go back and reconstruct that whole thing? So those are the four things we do. We onboard them. We assess and evaluate them. We manage the drift. And then we generate this.
[00:17:16] [SPEAKER_00] I call it the system of trust, the longitudinal system of trust. So you can go back and trace every flight path of every agent for both learning as well as for auditing.
[00:17:26] [SPEAKER_01] So you can see that record of what's happened. Is that in natural language? I mean, how does, you know, is the agents, consumer agents do, you can see the reasoning and then you can go back and look through it. Is it that sort of thing?
[00:17:48] [SPEAKER_00] Yeah, it's actually it's a it's sort of a graph. It's like a time series based graph that tells you all the various things that are coming into play. What models, what data, what policies, what constraints, what context. So it's a rich graph that actually is recording much like your video camera, your your autonomous cars, camera and navigation system records all the things that are going on.
[00:18:15] [SPEAKER_00] So it's a data structure that is can be thought of as a semantic graph, which can then generate tables and reports and can answer questions in natural language on top of it. But the data structure itself is one of our big innovations. We call it a semantic action layer, which is both used to drive the agent behavior as well as record the agent's behavior. That's one of our big innovations is the data layer.
[00:18:42] [SPEAKER_01] Yeah. Semantic action layer. Yeah, that's interesting. I've just been talking to people about vision action models and robots. So, yes. Yeah. And who is using this? I mean, in your who who in the enterprise, first of all, is worried about this? Is this the CEO, CTO?
[00:19:10] [SPEAKER_01] I mean, these issues are migrating up the the command chain. And then once it's implemented, is this an IT function? Is this a business unit function that where you have somebody that's that's checking the the the performance or, you know, who where is that human in the loop in the enterprise?
[00:19:38] [SPEAKER_00] Yeah, absolutely. So the three different layers, I'll break out the answer that way. At the topmost level, this is definitely a board and a CEO level issue now because you suddenly have a new risk surface in the enterprise. AI is a double edged sword. I mean, it's going to create amazing value, but also is creating new threat surfaces, new vulnerability surfaces.
[00:20:00] [SPEAKER_00] So the board is now getting into and saying, hey, if I'm, you know, introducing these digital workers, how do we make sure we are drug testing these properly? How do we make sure that each worker has the right employee handbook? How do we make sure we can do instant performance appraisal? And how do we make sure if they're drifting, is there a kill switch on these things? And then if we are challenged, do we have an audit log and evidence log? So that is at a board level. I like to say that AI is too important to be left to technologists.
[00:20:28] [SPEAKER_00] It's a business capability like email was as power was much more powerful than that. So that's at the top level. This is a top level issue at a CEO and board level. Then at the middle level, this is where the decision makers are is primarily CIO and the head of AI. They're ones who are struggling with deploying AI into production. One of the big things that we see is it's taking people a few hours to build an agent.
[00:20:56] [SPEAKER_00] It's taking them six to seven months to put that into production because they're not able to get the confidence. In fact, there was an MIT report that says 95% of agent projects are not able to move from pilots to production. And the reason is that you can write up an agent very soon, but risk and compliance need to make sure that it is following all the internal and external regulations.
[00:21:18] [SPEAKER_00] Security needs to make sure it is safe. Compliance needs to make sure that all the regulatory bodies in different states where they have different rules, it's going to behave properly. And then finance needs to make sure it's not going to bust the budget by going off the rails and taking up a lot of costs. So this process of really getting an agent production ready, that's what we do. We help compress the time almost like a spell checker for trust. Right.
[00:21:45] [SPEAKER_00] You can set up an agent very quick, but no one's people are compressing the time to build an agent. No one's compressing the time to put that into production and to control it when it's in the production. That's what Trustwise does with the control layer. And so that's at the second level. Most of our purchases are done by the combination of CIO, head of AI and head of responsible AI. Risk compliance and security also gets involved in it.
[00:22:13] [SPEAKER_00] But security looks at it as one of the three dimensions of it. Then at an operational level on a day to day basis, there are three different roles here. We call it the three lines of defense. The first role that's managing it is business and I.T. So they are looking at and saying, OK, how are my agents doing? What is their alignment to my policies? What is the output versus productivity? That's line of defense one. Then there is line of defense two, which is risk and compliance.
[00:22:39] [SPEAKER_00] Can I generate a report to make sure that I am on policy and it's not deviating from issues and creating new threat surfaces or risk surfaces? And third one is internal audit. Line of defense three is audit saying I need to make sure that these digital workers are behaving well. So if I'm audited by an external regulator or if there is an issue downstream, I can I can prove that the behavior was controlled.
[00:23:05] [SPEAKER_00] So that's kind of a tier thing, you know, board C-suite, primarily chief AI officer and CIO with support from the responsible AI officer and compliance group. And then finally, those three different roles can have windows and dials into the system through the AI control tower.
[00:23:24] [SPEAKER_01] One of those three roles or was it in the middle section you mentioned security, the cybersecurity team? Yeah. How and this is a little bit off topic, but I'm sort of obsessed with with this mythos thing from Anthropic, which is so powerful. You know, it's still the story isn't over, you know, the government.
[00:23:54] [SPEAKER_01] But that kind of power when you look at the curve of capability by agents and certainly mythos is at the top, but it's not going to be alone. The Chinese will be there. The Chinese will be there. Other foundation model companies will be there as that kind of capability spreads.
[00:24:20] [SPEAKER_01] How do you deal with that at the governance and trust and safety layer?
[00:24:27] [SPEAKER_00] I'm so glad you asked that because one of the things I've been telling boards and CEOs is that there are three things, three motions they have to manage around AI with the AI control tower. AI that you build, AI that you bring in by buying stuff, and AI that's going to come after your existing things, which is where the toast thing comes in. So we call it build, buy and protect. And so you've got to control all the three surfaces. What am I building with IT?
[00:24:55] [SPEAKER_00] What am I buying with vendors? And what am I protecting with my existing mainframe systems, my chatbots, my databases? And models like Mythos, we're just entering into a state where all hell is going to break loose, where the systems are going to start exploiting your existing assets. Even if you're not using AI, AI is going to come in and do penetration of your systems.
[00:25:21] [SPEAKER_00] So through the control tower, they have to be able to manage all three surfaces, and we help companies do that.
[00:25:27] [SPEAKER_01] And you are very focused on regulated industries. Is that right? I saw that Gartner named you Cool Vendor for Agenda AI in banking and investment services.
[00:25:42] [SPEAKER_00] Yes, we do. We call it high stakes industries. It has to be regulated. But my whole thesis was, you know, I've had the good fortune of building multiple companies. And one of the things I found is product and distribution is what is key to scale and build great companies. And I figured that banks have been doing model risk management for 60 years. And if I can go in and build a product which takes something non-deterministic like LLMs and agents and makes it deterministic,
[00:26:11] [SPEAKER_00] and if they can put that into production, I will build a sufficiently good product that I can take it to other industries that are not as mature. Oil and gas and chemicals and healthcare and others. So we started the company and spent the first three years working with banks building the product because we felt that's the highest bar we have to cross. And so our customers range today across banking and insurance, healthcare, even retail and consumer products.
[00:26:41] [SPEAKER_00] So Yum Brands, who've got, you know, 68,000 restaurants, they're looking at it for voice AI, voice ordering across Pizza Hut, KFC and Taco Bell. They need to make sure that the voice ordering agent is behaving properly at runtime. So we focus on high stakes deployments of AI where there is big value and big risk. And it tends to typically be in banking, insurance, healthcare. But at the same time, we also have other customers.
[00:27:09] [SPEAKER_00] And Hitachi actually is an investor in us. And Hitachi is OEM does as part of their AI stack. They have 650 business units that they are looking at OEMing trust-wise now as the control layer across all of Hitachi. So we're kind of building it like concentric circles. We first nailed bank and insurance. Then we added other industries. And now we are putting this as an OEM component into multiple platforms.
[00:27:36] [SPEAKER_01] From a system perspective, how does the runtime control technically integrate with existing orchestration? You know, because there is an orchestration layer in most of these stocks already.
[00:27:56] [SPEAKER_00] That's a great question. So we sit above the agent orchestration layer and below the user experience layer. So essentially, the AI control tower, we exist almost like a blip on the wire through a gateway. And we can be operated in one of four modes. We can be inline. So if you're orchestrating using a Microsoft agent orchestrator or Google or LangGraph, it doesn't matter to us.
[00:28:23] [SPEAKER_00] We are agnostic to agent frameworks and agent orchestrators. We are sort of like a probe or a proxy that's connected into those orchestrations. And that proxy can be invoked in one of four different modes. It could be live. So every piece of the traffic we are able to detect. It could be in a sidecar mode where we are sitting there and we only intervene when we see you drift. It could be in a batch mode where you can run all the data once every day, once every week. Or it can be in a simulation mode.
[00:28:53] [SPEAKER_00] Pre-production, you can simulate everything. So the actual runtime control has four different operating modes. Live mode, sidecar mode, batch mode, and simulation mode. And it is agnostic to agents, agnostic to agent frameworks, agnostic to cloud, agnostic to models. And we have deployed this across, again, like I said, in very tough banking environments. We have deployed this on Azure, on AWS, on GCP, on Google.
[00:29:23] [SPEAKER_00] And we have also deployed it on on-prem, on Dell in a partnership with NVIDIA. So we are agnostic. This is the part that's really important is this control layer has to be able to sit across agents coming from your teams that have built agent using maybe LandGraph. ServiceNow has given you some agents. Microsoft Copilot has given you some. Cloud has given you some. So just like with security, each of these players, they have their own control plane.
[00:29:51] [SPEAKER_00] But the company needs a meta control plane across all of these. And that's where we sit as a vendor agnostic layer. That's why we call this a NatWest Bank control tower, a US bank control tower, a Yum Brands control tower. And that sits across all these different vendors. When it detects in the proxy time, we have a collection of 12 very small, high precision, low latency models.
[00:30:21] [SPEAKER_00] And we have 15 scales and orchestrators. And then we have also a collection of the whole IP that goes under it. But when we connect to it within 10 seconds to 300 milliseconds, we're able to do those evaluations at runtime. So as your agent is running with a probe, depending on the type of workload, within 10 to 300 milliseconds, we are able to steer the output of the agent.
[00:30:46] [SPEAKER_01] And first of all, this platform, you guys call it Optimize colon AI. Is that right?
[00:30:55] [SPEAKER_00] Yeah, that was the first release of the product. Now we call it Harmony AI. So Optimize AI was the first release, which was primarily on generative AI, where you can take LLM based system. Now the release we have is called Harmony AI, because now we have to have multiple agents living in Harmony across multiple vendors. So the current product release name is Harmony AI, but the product category name is the AI control tower.
[00:31:20] [SPEAKER_01] You gave an example of an agent that gets stuck in a loop and is burning tokens, which of course is money. And there's been a lot of talk in the last couple of months as these systems move out of pilot mode into production that people are discovering that this can quickly get very expensive.
[00:31:47] [SPEAKER_01] Agents talking to each other and that sort of thing. How does your system control that or surface that or because you do talk about a substantial cost reduction when using Harmony AI?
[00:32:08] [SPEAKER_00] Yeah, we have demonstrated as much as 83% reduction in costs as we improved safety by 40% and latency by another 60%. So we have a lot of case studies where we have proven this. The core point to this is when we move from generative AI to agentic AI, the number of tokens, even though token costs are coming down, the volume of token consumption has gone up massively.
[00:32:32] [SPEAKER_00] So one agent action can, one input into an agent can trigger between 20 to 50 actions and can consume between 20x to 40x more tokens than you would with a generative AI system of two years ago. So suddenly you have one input driving maybe 50 different actions and driving 20 to 50 times or more token consumption. So it becomes incredibly important to get efficient with these things.
[00:33:02] [SPEAKER_00] The way we help companies manage the tokenomics as they call it is, number one, when we assess, so when you bring in an AI system, the first thing we do is we classify it. So we use a framework from the Responsible AI Institute called TrustX. And Responsible AI Institute is a 10-year-old nonprofit. It's the oldest and the largest nonprofit, which lets you understand what kind of a system are you dealing with? Is it a golf cart? Is it an SUV? Or is it a tank?
[00:33:32] [SPEAKER_00] What are you building? So the first thing we do is we assess what is the expected performance of a golf cart or SUV from a fuel consumption point of view. Second thing we do is we then say, given that you've got an SUV, it's not a golf cart, it's not a tank, here are the policies and controls you should put in place so that you're evaluating the right behavior. You're not excessively loading it up with too many things or you're not being short-sighted.
[00:33:58] [SPEAKER_00] And third is we simulate it and we run the performance runs in our sim before you put that into production. And we will then tell you across the whole pipeline what tweaks to make. What should be the chunk size? What should be the number of chunks you should put in? What models should we use? What endpoints should we bring in?
[00:34:18] [SPEAKER_00] And it's almost like a Cisco router where you're figuring out how for that given workload, what's that optimal configuration that gives you the right safety, the right compliance at the lowest cost. And this is one of the core patterns in IP. When I said our engine does 10 seconds to 300 milliseconds, that's what it is running and making those tradeoffs against. So it's a three-stage process. We classify it to understand what the tokenomics should look like.
[00:34:45] [SPEAKER_00] We simulate it with thread packs and performance packs and see where the pipeline should be tuned for the most optimal performance. And then at runtime, after it is running, we actually use agents to discover which agents are drifting, which are actually maybe too tightly controlled that they're using too much tokens. And we will generate reports after it is running to say, hey, you can save another $8 million across your 1800 agents if you make these changes. So it's all three stages.
[00:35:15] [SPEAKER_00] When you bring it on, when before you put that into production and after production.
[00:35:19] [SPEAKER_01] Yeah. And what's your feeling is there's been, I mean, it's been a little bit now, but there was a lot of people writing about how these agentic systems, which are supposed to do the work of employees are, you know, multiples more expensive than the employees.
[00:35:44] [SPEAKER_01] And so does this kind of governance bring that down in line with human labor?
[00:35:57] [SPEAKER_00] Absolutely. So I think, so I'll answer it in three parts. Number one, we are just getting going with this stuff. You will have different classes of workers, just like you have in a company. You will have workers who are low cognitive overloads. I mean, there are people who are at the gate who are helping you check in and check out versus you have engineering designing, you know, three dimensional models. So you will have different classes of workers with different tokenomics associated with it and different economics associated with it. That's number one.
[00:36:27] [SPEAKER_00] Number two is these workers today, they're designed with one model connected to that one worker, and it is not dynamically shifting the models. So you can see that in the past when you connected the phone jack into it and you left it. Today, when you look at a cell phone, it is dynamically switching the towers to get you the right performance. So that is the second part. Model routing and model selection underneath these agents is going to come up as a way to start managing the optimal performance of tokenomics.
[00:36:56] [SPEAKER_00] And third and final one is there will be a whole new collection of analytics and insights, which is what we provide to start surfacing where you are drifting, where there is opportunities for optimization. And in many cases, you can even build agents that will autonomically start adjusting those models for you. We're not there yet, but we have demonstrations where we have proven that that can be done. So all three put together is going to drive, I think, where this thing is going.
[00:37:25] [SPEAKER_00] But I can definitely see for certain agents, the cost of operating that agent is going to equal, if not exceed the cost of an employee per year. It's definitely going to happen.
[00:37:35] [SPEAKER_01] So, this market, this new market that you are helping create or new category, what is the competitive landscape? I mean, it's not very crowded yet, the building these trust and governance frameworks. But there are other vendors and there are internal policy engines.
[00:38:05] [SPEAKER_01] How do you look at the competitive landscape?
[00:38:10] [SPEAKER_00] Absolutely. So, cyber trust, we call this a cyber trust market versus cyber security. Cyber trust is going to be as big, if not a bigger market than cyber security was. And we're just getting going. By some estimates, it is supposed to be worth $80 billion by 2030. And it is probably $2 billion this year. And so, that's one. It's a large emerging market.
[00:38:32] [SPEAKER_00] It's a white space that a lot of companies, governance companies, security companies, observability companies, they will all try and get into it to get a piece of it. That's number one. Number two is, this space and this problem is going to be fundamental to how effectively you will put AI to production or not. My thesis for building the whole company is, intelligence without control is not deployable. And that's why 95% of projects are failing from going into production.
[00:39:02] [SPEAKER_00] I can build you the most, you know, biggest model that can do wonderful things, but if it doesn't align to your policies and if you can't control how it is running, no one's going to allow you to put that into production. That, to me, is the biggest risk, is the issue is not building an agent. The issue is controlling an agent. So, where the market is going to, where the problem is, is next enterprise AI problem is not whether models are smart enough.
[00:39:28] [SPEAKER_00] It's whether autonomous AI can be controlled while it is operating. And the people who are approaching it from a competitive landscape point of view, there are four different ways that this is being handled. One is security companies approaching into it saying, I'll secure the behavior. Second is governance companies. These are mostly declarative governance companies who are saying, I'll make it into a runtime governance. Third are observability companies who say, I'll monitor it.
[00:39:55] [SPEAKER_00] And then fourth are people trying to roll their own internally. Like you said, taking the role of existing policy engines. What no one's doing is bringing all those three postures together. We call this trust posture management. You need it to be secure and you need to be compliant and you need it to be cost and carbon efficient. And so we have a lot of competitors, but they're only solving one of these pillars.
[00:40:19] [SPEAKER_00] There are no one solving as well as I know all the pillars put together and across all vendors and all platforms. We, the fact that we are agnostic to cloud and models and agent frameworks makes us very attractive. I'm seeing this in my sales cycle. Our sales are going nuts right now. Large enterprises that used to take nine months to make decisions. Within one meeting, they're saying, okay, let's start thinking about how do we put this into POC? Because this is becoming a burning platform now.
[00:40:48] [SPEAKER_01] That's fascinating how fast this is moving. I mean, it must be, I mean, it's a new game for CEOs and C-suites.
[00:41:02] [SPEAKER_00] I mean, my head hurts. I mean, being, Craig, being in this market, being in this industry for 12 years, I'd like to think I know a little bit about AI. My goodness. I mean, I myself, you know, as someone who knows the space, I've has collected the scar tissue for the last 12 years. I can't keep up with this stuff. And that's where I think keeping a real focus on outcomes and not on models. At the end of the day, companies are not buying AI. They're buying business outcomes.
[00:41:29] [SPEAKER_00] And whoever can get you to business outcomes with the lowest amount of risk and the highest amount of efficiency is the one that's going to win. So I keep telling my guys, we are not a software company. We are a company that enables businesses to unlock the potential of AI in a risk mitigated manner. We are, if you want to call it that way, we are an outcome enabling company with AI.
[00:41:54] [SPEAKER_01] From a regulatory and risk standpoint, how well does trust-wise map its controls and evidence generation to specific regimes you care about? And I'm thinking of, you know, SEC or FINRA for trading adjacent use cases or HIPAA in healthcare, the EU AI Act. Yeah.
[00:42:22] [SPEAKER_01] How do you stay up to date on that? And how do you, yeah. Yeah.
[00:42:32] [SPEAKER_00] That's a really important question because as a part of the platform, we provide over 1100 controls spanning 17 different regulations, pre-built policies that are verified. So things like NIST AIRMF, EU AI Act, OWASP, HIPAA, SR 11.7, UK FCA Consumer Duty. So we provide out of the box a library of these runtime policies.
[00:43:00] [SPEAKER_00] Almost think of it like Norden Antivirus packages that you keep adding these policies and controls. And every 30 days we update those libraries. So that's a big accelerator for customers. Most customers don't even know how to go about building these policies for runtime. So we provide a library as an accelerator. Plus, we provide a tool where the customers can add their own policy rapidly.
[00:43:21] [SPEAKER_00] And we have services partners like Accenture and KPMG and Hitachi Digital Services that can go in and do services work for you to create new policies for it. So that's a key accelerator in addition to our evaluation engine for trust. The library of assets that we provide, these 1100 plus controls, and then providing this, you know, runtime control generator tool is a good part of our value problem.
[00:43:49] [SPEAKER_01] All these layers that are stocking up. So there's just most recently the agentic layer, the optimization layer, orchestration layer, now a governance layer. Does this slow down processes? I mean, granted, these are processes that move at the speed of electrons. Yes, yes.
[00:44:19] [SPEAKER_00] Very good question. I think one of the things that people don't understand much yet is that AI is not an application. AI is an actor. You are giving birth to an actor who has got the capability to sign into websites and into databases. An app. Last 75 years in IT, we have built applications. We didn't build actors. That app stays the same. It waits for you to come give it the input. It doesn't evolve. It's rules based.
[00:44:50] [SPEAKER_00] These are pattern based, intelligence based, data driven systems that are doing things on their own. They're evaluating. That same actor will give you a different action or a different answer tomorrow than it does today. So as a result, there is a whole new enterprise stack that needs to be built up. And that's what we are a key layer in there, the runtime control layer. So broadly, there is the compute layer at the bottom. There is the data layer. Then there is the model orchestration layer.
[00:45:19] [SPEAKER_00] Call it model and agent orchestration layer. And then there is the runtime control layer. And then the top one is the user experience or the machine experience layer. Yeah. Because a lot of the consumers of this are also the consumer of our product, frankly, is not humans. It's agents. I expect in three years, 90 percent of the people who will be using the AI control tower will be agents using it, not humans. Because I don't know if you track this.
[00:45:46] [SPEAKER_00] But last month, the first time ever, the traffic on the internet, agent traffic exceeded human traffic.
[00:45:53] [SPEAKER_01] Oh, is that right? I didn't see that.
[00:45:56] [SPEAKER_00] Yep, absolutely. It's a milestone shift. I was going around telling my family, they were looking at me crazy. I'm like, this is a very big deal. Just like when on the AT&T network, data traffic exceeded voice traffic. We are old enough to know we were there back then. But now the same way. So our product actually is built for both humans and agents to consume. Because at the end of the day, this stack will be operated more by agents at the speed of electrons than it will be by humans.
[00:46:26] [SPEAKER_00] And there's a whole category there called AX versus UX, agent experience versus user experience.
[00:46:32] [SPEAKER_01] It's amazing. I've said this a few times on the podcast, but I was talking to somebody in the... drug discovery space. And they were saying that, you know, they have a platform that prepares the documents to send to the FDA. And it's an AI that generates this, you know, tome of writing.
[00:47:02] [SPEAKER_01] It's sent to the FDA and the FDA feeds it into an AI to read it. And they were saying that soon they're just going to drop that middle layer and the AI will send data directly to the FDA. And it'll all take place outside of natural language.
[00:47:25] [SPEAKER_00] Absolutely. Multi-agent, multi-enterprise workflows is truly where it's headed. That's why I said in the beginning that OpenClaw is 100 or maybe 1,000 times more significant innovation than ChatGPT was. The problem with OpenClaw is currently there is no control and there is no policy alignment. So they look like a collection of mad barking dogs.
[00:47:47] [SPEAKER_00] But if you start putting the steering wheels on them and putting the policies, you can then guide them to do some amazing things for pennies an hour and automate and come up with new things that we have not even dreamed of yet. So I do believe this is one of those. I had given a TED Talk on this like seven years ago when I had said the death of Homo sapiens and the beginning of Homo digitalis. So this is the beginning of a whole new species that we have summoned.
[00:48:17] [SPEAKER_00] And I don't, people have really kind of gotten their heads around it yet, but I'm seeing glimpses of it. And what is important is when you have summoned this, I call this AI is not artificial intelligence anymore. It's AI in some cases is alien intelligence. And if you have alien intelligence, how do you control it? How do you make sure it is aligned to your intent and your business goals and societal goals? That's the problem we are solving at TrustWise.
[00:48:43] [SPEAKER_01] And how do you do that at scale? And absolutely, I can see that this trust in governments is becoming increasingly critical as these systems expand. I mean, presumably we'll have a sort of underlying fabric in the economy that's agentic. And the entire economy is going to be depending on that.
[00:49:11] [SPEAKER_01] And you want to be sure that it's aligned. That was one of my questions on alignment. How do you, is that align a system with your values or your goals? Is that done through prompting? Is it done through fine tuning?
[00:49:34] [SPEAKER_00] Great question. So I'll answer it two ways. One is, first of all, prompting is rapidly becoming passe. Prompting is being replaced by loops. Okay, so loops is the new thing. Where these systems are going to now set up their own experiment. They're going to take their actions. They'll review it. So that's how this whole autonomy, autonomous behavior, where these systems are running for hours and now soon days. So that's one shift.
[00:50:00] [SPEAKER_00] And you've got to make sure that it is aligned as it's doing these things autonomously for a long time. The second way I will answer that is alignment exists in our framework. And this is configurable. We have six layers of alignment. So when an AI is acting, it has to align at six different levels. One is at a global rights level. So UN Human Rights Charter. It has to have values at that level you have to align to. Second is national level.
[00:50:27] [SPEAKER_00] Singapore AI Act versus U.S. versus Saudi Arabia. That's the second layer. Third is industry layer alignment. FINRA versus HIPAA. It has to have that. Fourth is company level alignment. What are the company's values? American Express, for example, they have started putting corporate values into their agents. So that's fifth is business unit and business workload level alignment. Claims processing agent versus customer support agent versus HR travel management.
[00:50:57] [SPEAKER_00] They have a different set of alignment. And last but not the least, customer SLA and customer engagement alignment. So when we do that evaluation alignment between 10 seconds to 300 milliseconds, we are actually aligning all those layers of the stack and ensuring that you are steering the actions at runtime. That's a really important problem. And any of those layers slip, you're going to have a catastrophe on your hand.
[00:51:21] [SPEAKER_00] And if you can't prove later on what flight paths did you take that put like a toothpick through the sandwich, as I call it, can you prove the toothpick through the sandwich? And can you prove to a jury that this actually did something that was aligned to all those layers? And that's the heart of our technology.
[00:51:39] [SPEAKER_01] You said prompting is becoming passé because of these looping transformer architectures. But still, the transformers need to be looking up.
[00:51:56] [SPEAKER_00] Yeah, they need to be looking up values and the charter. Absolutely.
[00:51:59] [SPEAKER_01] So is that a RAG system? If it's not a prompt, how does the system get access? Yes.
[00:52:08] [SPEAKER_00] Now you're getting into the real IP now. I'll have to kill you if I tell you. But this is where I believe the current technologies are failing. This is where the new data layer that we have built is important. With the semantic action control layer, that's where we believe. So I'll give you an example. The whole thing comes down to what is allowable and do you have a data structure that tells you what permission do you have at runtime?
[00:52:34] [SPEAKER_00] That's the heart of our data technology, which is it's not enough to say I have a knowledge graph. Knowledge graph just tells you relationship. Okay, you've got a Mario Brothers game. You've got trees that are connected to ground. Ground has water. That's just a relationship. Then you have context graph that says, okay, in this, the apple is connected to a tree and then Mario can run through the water and run through the gap.
[00:52:58] [SPEAKER_00] Then there is world models, which is what Jan Lekun is all about, which is saying an apple typically falls to the ground because it's a thing called gravity. None of these are enough to help you control a runtime behavior on agent. There is a layer beyond that that we have invented and innovated called the semantic action layer, which is what are the allowable action paths for that agent at runtime that aligns to all those seven layers? That's what I've been working on for three years.
[00:53:26] [SPEAKER_00] And that's what we've implemented into the product is what is the permission pathway that the agent is allowed to execute on? And that's where we get to make AI deterministic. That's where we get to make agents provable. And that's where we get to make agents super fast. It's like if you're running an autonomous car, every 10 meters, the car is recomputing the map. That's how people are building agents today.
[00:53:54] [SPEAKER_00] We have built a whole different thing where we've actually built a Google map which tells you what is allowable. And then you can run through it very fast by reducing cost, reducing latency and creating provable. That's what we have patented and that's what we have implemented.
[00:54:09] [SPEAKER_01] Yeah. And so if you're changing geographical regions or government systems, it's sort of unplug one foundational document, plug in some others. Exactly. I'm going to ask, this will be my last question. We're running out of time. But this is the sort of pie in the sky's question.
[00:54:38] [SPEAKER_01] I worked for a little bit with a guy at Boston Consulting, a brilliant guy. He's now moved on. He's the CEO of an insurance company now.
[00:54:53] [SPEAKER_01] But he was talking about, you know, there's 5,000 years of human moral teachings in the great books of religion and philosophy.
[00:55:12] [SPEAKER_01] And is there a way to distill that knowledge, de-conflict that knowledge to come up with a guiding set of values for AI going forward? What do you think of that? Or do you think that is already baked into the foundation models?
[00:55:39] [SPEAKER_00] It is such a wonderful question. Now, if you really look at what these foundation models have done, they have absorbed all the knowledge. Of course, it's only Western knowledge. It doesn't include other countries and stuff. But, you know, it's imperfect, but it still is a snapshot of humanity. If an alien descended on the planet, they will say at least the Western world, these are the values, the ethos and more is going back all the way to Greeks. And even before that, that's what these models have. What we've actually done, I actually have a prompt.
[00:56:07] [SPEAKER_00] I'll send you that, which can tell these models, explain to me what humanity has said to you about values and religions and stuff. And it will tell you what it is. I've actually written that in a dialogue with the foundation model. And it is beautiful. It tells you about love, tells you about sacrifice, tells you about values and what we talk about and what we act are different things. That's what the models have learned.
[00:56:31] [SPEAKER_00] Now, you could distill those values and you can implement that as one of the core alignment pieces in your AI system. But companies are already doing that by starting with corporate values. They're saying, leave the humanities values. American Express is a good example. They are starting with saying, here are my corporate values. I want every agent to follow this. So I do believe we are stepping into a whole new area of taking those and implementing it and also coming up with new values.
[00:56:57] [SPEAKER_01] Yeah, if you could send me that prompt, I'd love it. I'll play around with it. Because one of the questions is identifying areas of conflict between different value systems. And as you said, the internet and the current models are largely trained on Western thought.
[00:57:18] [SPEAKER_01] But I would imagine in China, and I haven't spoken to anybody about this, but there must be a lot of research into distilling, you know, 3,000 Chinese, say 5,000 years of Chinese thought. But, yeah.
[00:57:37] [SPEAKER_00] In fact, given your background with, you know, New York Times, Wall Street, you could actually do a whole episode yourself, Craig, just on interview with the foundation model, where the foundation model using my prompt will tell you what it has understood humanity to be. Unmonished version of it. Unmonished version of it.
[00:58:25] [SPEAKER_01] No, no, no, no, no. So can you talk about that new product?
[00:58:28] [SPEAKER_00] Yeah. So basically, this is the next release of the Harmony AI Control Tower, where we will have two classes of agents in it. One is Guardian agents. Second is Genesis agents. So the real power, everything they're doing so far in the control tower is mostly to prevent bad things from happening. So that's what the Guardian agents are doing. But the real unlock is making the AI discover the unknown unknowns. So when I was running IBM Watson, I met this chess grandmaster who had lost to Deep Blue.
[00:58:57] [SPEAKER_00] And I said, what happened? He said, you know, most grandmasters can look 22 to 24 moves deep. This machine was looking 95 moves deep. We didn't even know there was a move like that. So what we are launching is the first prototype version of an open claw-based system where a group of agents can act as those machines that can go in and look 95 moves deep in domains like revenue leakage into fraud. So the next part that we are going after is
[00:59:26] [SPEAKER_00] once you got the agent settled down, how can you now unlock the value through a whole new class? People call this AGI, but verticalized AGI where the machine can start selling the CFO things where they are losing revenues, where the deals can be better by answering questions that even the best CFO on the planet doesn't know how to answer. So that's the beginning of our journey towards adding two classes of agents within the control tower, guardian agents and genesis agents.
[00:59:54] [SPEAKER_01] And the genesis agent is the one that's looking.
[00:59:57] [SPEAKER_00] 95 moves deep, a thousand moves deep. And it's surfacing hypothesis. So think of it as a hypothesis generating machine. So we call it a beneficial hallucination rather than a harmful hallucination. I believe hallucination is both a feature and a bug. So once you have the open claw agents settle down, we are now using those agents to create beneficial hallucinations and insights. And maybe only two out of 10 is right, but those two could be game-changing for your business.
[01:00:26] [SPEAKER_00] And we have just finished a couple of pilots on that with banks and they have come back for more. They are like, this is game-changing for us. Wow. So that's the view of what's coming.
[01:00:36] [SPEAKER_01] That is powerful.
[01:00:37] [SPEAKER_00] Yeah. I think this is the real exciting part, to be honest. This is HCI. This is the beginning of, we call it autonomous finance. And this is the beginning of AGI into finance.
[01:00:50] [SPEAKER_01] Yeah. And the underlying models, as we said with Mythos, are going to be so powerful that they'll be able to discover things that a human could not.
[01:01:06] [SPEAKER_00] See, this is the important part is the big players like OpenAI and Anthropic can't get to solving this because much of the context and the policies belong to the company. Yeah. So the real data to make that AGI happen actually belongs to the enterprises, not the model providers. And that's why I don't buy the Silicon Valley frou-frou stuff about, oh, we're going to create AGI. How will you create it when you don't have the policies to align what good outcome looks like?
[01:01:35] [SPEAKER_00] And that's why enterprises that have data and policies, they are going to be massively transformed in the coming three to five years.
[01:01:42] [SPEAKER_02] Thank you.

