Steadcast
Eye on AI cover art
Eye on AI

AI Agents Fixing Your IT Before You Even Know Something Broke | Erhan Giral & Ryan Manning, BMC Helix

August 3, 202659 min · 9,211 words

Show notes

Most enterprise IT teams spend the majority of their time fighting the same fires repeatedly. BMC Helix is building the AI system that handles those fires automatically, detecting anomalies, tracing root cause through millions of asset relationships, generating remediation plans, and learning from every incident it resolves.

Transcript

Automation Impact

0:00What's that going to do to IT staff? It sounds like it's going to get increasingly automated. That space is changing pretty quickly and that workflow is changing agentic AI pretty dramatically. Right now, AI, if you think about it, AI is only learning through someone else's description, right? People will start exposing these models to the real world more and more so that they can create, they can gain that first-person view of things. Once we detect something, we always ask the data, why, why, why, until we get to a point where we can take an actionable step to mitigate or perhaps even

Guest Introduction

0:32remedy that issue. Why don't you introduce yourself to listeners? My name is Arhan. I run the AI office for BMC Helix now. I work with Ryan on applying AI to various different problems in service ops space. Service ops is essentially, you can think of that as any digital business these days is really a conglomeration of IT services, you can imagine. And service ops is essentially an attempt to automate operational activities around these

1:08services as much as possible. So that we look at enterprises as combination of software, infrastructure, and human beings operating on these assets. And we provide various products and nowadays agents to automate different workflows around these services. My personal background is in monitoring space. I've done early work on monitoring applications and infrastructure and then combinations

1:44of them, which means you have to process a lot of machine data. You have to be able to reason about high-throughput streams of data that these data centers and these machines are constantly emitting.

Helix and IT Service Management

1:59So yeah, that's what I do in Helix. We're going to talk about Helix and IT service management. Can you begin one of you by talking about how the landscape has changed with the introduction of AI agents, what the process for service management was before, and then what it's migrated to today?

2:34So yeah, you've got the recommendation of you by getting some reports that you'd like. Bye. Bye. Bye. Bye.

2:44Bye. Bye. Bye. Bye. Bye. Bye. Bye. Bye. Bye. Bye.

2:57Bye. Bye. Bye. Bye.

JIRA and Service Management

3:02Bye.

JIRA and Service Management

3:03Bye. Yeah. And, and, you know, just thinking back in the old days, it was JIRA, right? You would, you'd open up a JIRA ticket. It is, it is, does your solution work alongside JIRA, replace it? Just for people that aren't that familiar with, yeah. Yeah.

3:34Yeah.

4:04right and and and you guys uh we'll give the background of bmc first you were talking about that uh before we started

Service Management Discipline

4:50yeah yeah so i mean service managers the discipline is um is complicated the the explanation of what they do is pretty simple it's the front door uh to it so if you know i'm a new employee trying to get my laptop provisioned or if my if i'm having an issue um with an application um i might go to a portal uh send an email call the service desk and try to get that um remediated and um that that

5:22that space is changing pretty quickly and that quit that workflow is changing um with a agentic ad pretty dramatically

Market and Competition

5:52yeah no that's um um jira does have it so our our competitors you know go beyond service now um into atlassian into fresh works um the market is a lot of players the interesting thing about the service management market is that unlike the cri market to broaden it out for a more exceptionally difficult to general audience uh describe what it service management is and how not just from a product critical it is to enterprise and so in the high enterprise where we operate um we'll say the

6:29fortune 2000 um you submit a ticket you're either submitting it to service now or to bmc helix down market um jira is a very popular tool fresh works a very popular tool i volunteer a few others yeah yeah i like to think about bmc in four eras uh era one remedy rules um a lot of folks

7:05know remedy um that was an acquisition that bmc made years ago uh and was the service now and service management before service now it was the on-premise uh service now um era two uh about 2010 the music began to slow a little bit um lost our way a little bit a little bit in terms of you know uh hopping on the cloud uh train um and then the company ended up getting bought by bain and taken private um sort of

7:36continued the same mission um yeah and there were three uh kkr came in and said hey service now has one competitor it's bmc helix but what if we took a different approach what if we brought in a technical team in the operations of those companies just recently turned around and see it's kind of the critical um you know ai first um and that that journey you know took us to close to where we are

8:13today um where we you know went through the the natural progression of uh the heartaches for our customers and for ourselves and and rebuilding a platform and still servicing some of the biggest companies in the world to now being able to innovate and separate ourselves from the competition with the innovation um that we got access or permission to build um through kk arizona ownership

Service Management Requests

8:44um yeah yeah and there's a service management being the front door to it can encompass a lot of different requests that need to get fulfilled some of them are very simple you know i need a new iphone um some of them are hey the quoting tool is down it's the end of a quarter um we're in trouble and there's all these workflows that sit behind that request some of them more complex um when those more complex general model issues come in amazing capabilities in the human

9:20language the front door team talks to the engine room i think everyone you know keeping those services up to date what we've been trying to do is and there's a there's a process to remediate that issue and so what we did in that transformation i talked about is we brought those two disciplines together the various sorts that these machines naturally emit category service ops um because when you um because when you end up sharing data um because when you want to reason about that massive graph ryan was talking about communication you have to be able to quickly process all the telemetric

9:52data you have to worry about lots of outages taking too long to be able to restore that outage understand that graph and those relationships like less regression to the same thing happen natively uh and in it uh it's very common that you have to constantly deal with deep relationships because something that runs on a computer turns out it's actually not a real computer it's a visual computer that is resourced by a physical computer somewhere else so this is all these resource dependencies resource and transactional dependencies in the fabric that you have to understand uh before you can be

10:26um you know you can produce anything actionable and useful because on the on the operation side and what we want to uh complicated things so uh you know on the service management side a lot of disruption there's a lot any language-based workflow is right for disruption so uh when i used to have to go to a portal other agents that other people can take action and that does require um to make to be able to be able to do that in a reliable sense um you have to create uh repeatable predictable

10:59infrastructure from the machine and from these models and there is so many relationships between uh investment in that regard to it's just amazing we have a customer with only 200 000 assets in there together and um and by cmdb where they store all their assets can you walk us through 18 million in relationships a use case or a case study and so go into that space it's not like i was just plugging lm and and oran hopefully you can yeah so uh and you so just like you know ryan said from

Root Cause Analysis

11:30from outside in it's actually very simple so you first need to understand around that hey do i have a problem to get to root cause why is that happening should i mobilize anything at all like do i have a problem like awareness and identifications of problems is one use case you know we constantly uh deal with and that requires our agents to stream this telemetric information this can be your machine logs time series information alarm information that they emit uh so something needs to be there and constantly

12:02monitor these things and um we have a funnel that deals with that you know let's take the observable data and run that through various machine learning models so that we can understand oh you know that spike that you see right now at 8 30 a.m on a monday on the bank's atm network is actually kind of unusual that that that you don't expect that spike to appear at 8 30 perhaps maybe it's supposed to come in later uh so something needs to essentially constantly reason about hey you know what's unusual or uh or or given given the time we are in or given the workload we are dealing with

12:38given the circumstances in general we are so that's one part of the big use case we try to do and of course once you detect an anomaly once you detect and uh available the problem you then first ask okay so why you know why what is the root cause of this issue um you know why why am i why is this router so busy that you know where it where it should be you know close to idle state for instance so then that we call that the root cause analysis so essentially something that needs to dig into all that telemetry data dig into all that observability data to say ah it's you know it's actually it's not that that's that's a

13:13symptom uh what's really going on is you know such and such file system or such and such uh you know network queue such and you know in broker somewhere is now full and you know dropping the traffic or what have you so uh you know you always ask uh you know once we detect something we always ask the data why you know why why why until we until we get to a point where you can take an actionable step to mitigate or remedy that or perhaps even remedy that issue um and uh while doing that of

13:47course you know data centers and it is very target rich all kinds of stuff happens all the time uh so you need to be able to do this um uh by sifting away all the noise all the background noise that typically happens in in that environment that's uh it's like the you know there's always this cosmic background noise that you have to essentially first detect and then push away to then focus on the real signals in the data uh to understand what the root cause is uh and then um and then we also do impact analysis

14:21like okay so is this a big problem because sometimes technically challenging problems are just that you know no one no one cares about maybe it's a qa environment maybe it's a staging environment that's being you know modified at the moment uh so essentially from economic point economist point of view we always ask okay we have a problem but is this an impactful problem should we mobilize human beings should we spend more resources in mitigating this issue uh so that so our system constantly makes these decisions in the background by uh looking at real uh real-time status of systems while also constantly

14:59comparing um the the state to its past self and um we also make you know plans of action uh based on um plans of actions that were executed in the past by human beings or or or or by by bots um and this well first of all this is all being done um within an agentic layer it's not uh it and with reasoning models to look

15:29at the problems and try and figure out it's not uh kicking uh questions out to the the human team uh who then sit around and try and figure it out it's it's doing this internally right that's right so some parts of what i just described that is actually taking that uh you know gigabytes of gigabytes of monitoring data and then reducing and then uh essentially pushing that noise away that's typically uh we

15:59employ proprietary technologies to uh to to to to to do that because uh there's so much data to process uh we couldn't really hope to expose all of that to the generative models all the time um so we take all that uh sparse uh data uh and run through various machine learning and uh statistical uh analyzers first uh to to to to basically find the needles in the haystack if you will but once those needles are found

16:31uh they are collated combined correlated causally analyzed uh and then the uh the llm you know the fine uh uh llm is uh given a pretty comprehensive causally uh uh described uh like a like a like a a patient chart like a like a set of x-rays of the system and then uh we asked that we asked llm okay um i mean we asked various questions to llm but most importantly uh you know we asked the question why

17:05and you know what needs to be done and that level is completely agentic just like you said um and there uh we use um uh reasoning models uh just like you said maybe with a twist uh we because we want our models to reason just like one of the employees of our customers um because like for instance if you go to an open ai model or uh a generic of the shelf model um they'll always give you plausible uh responses you know accurate responses based on the documentation based on what they know about the

17:37space but it's often the case that when you go to a big organization the rules of engagement around these systems are very different meaning you can say oh you know go to this file and make these edits and everything should be fine uh but then the that banks uh knock you know network cooperation uh personnel will probably say something like well you know that sounds plausible but that's not how we do things you know i need to first you know run these additional tests maybe i need to talk to some other team get their approval uh you know i need to make sure you know there's a backup procedure

18:13i need to estimate this there's there's lots of lots of things that they um need to uh account for uh and then that's often specific to the uh enterprise to the to the uh to the to the day they work so what we do is we take our uh we take the uh reasoning capabilities of these models and we fine-tune them so that they um they can they can reason just like that employee of that particular enterprise uh so we found that um when you tune these models especially the planning aspects of these models

18:46um around the subject matter experts practices uh they become extremely extremely um um capable uh they they really produce actionable uh actionable and relatable results for the users um so just to point to your um question about uh reasoning yes we we rely on reasoning uh but uh we we try to guide we try to guide that reasoning towards uh economic ways of solving that problem that's another thing great so a lot of times when you ask uh complex it questions to these models

19:19um you know they'll they'll come up with you know plausible you know possible plans but then is the is the plan actionable essentially if the solution is you know what uh tear down everything rebuild everything this is a rewrite situation you know then the customers won't be happy you know they'll they'll just say that's too expensive i can't really do that um so then we we train our models and we reward them in ways uh such that uh more economical uh more uh more expedient ways of uh lower risk

19:50solutions are preferred over others and and what so there's a lot goes into the guidance uh guidance of that reasoning but at at the court just like you said we do rely on uh llm's reasoning capabilities uh to figure figure figure things out to do the orchestration especially yeah uh and and so that that whole reasoning uh workflow that's is that one module i mean we were talking uh before we started

Modular Architecture

20:17recording that you guys uh build uh your modular so you have these microservices that work together uh that presumably then you can uh configure or or pick and choose among them uh depending on the use case uh is all of that reasoning one module yeah yeah so uh we follow the similar philosophy of uh divide

20:48and conquer you know create good encapsulations around uh around capabilities um so we follow an architecture pattern called mixture of experts yeah what this is is it allows you to um you know you know these pre-trained models are trained on piles of piles of publicly available textual data i mean if you're talking about language models um and uh what we found was actually what we not we didn't find this but uh meta sort of led the way into this architecture uh in our in our adoption uh they um they described uh

21:24they had this wonderful people paper uh to talk uh where they talked about you know how how these models can be further tuned for different purposes while encapsulating that training in a specific set of weights and biases uh called experts uh so what that is is you take one of the uh open weight models um open weight model means a model that be you can basically download your computer and run as much as you want based on of course depending on the licensing agreements um so what you do is you take one of these

21:59open source models and then you identify a problem uh data set uh in a reward function so you essentially understand understand you you basically device a clever way of signaling the network hey you know uh here is uh if when you when you give me a good response here you know i'll give you uh cookies uh if you if you don't uh if you if you if you fail you know i'll you know you'll you'll get the stick so you basically train train these models on additional data sets that allow you to focus on that problem uh and then

22:33you capture these as what we call experts so experts are just like you're saying they they're mini uh i mean that's i wouldn't say they're services but they're encapsulations of that training that they're the artifact of that training cycle uh and um within this mixture of expert architecture you also specify gate activation uh values during your training which means the model then knows when when it's prompted with a question whether to activate that training set or not like for instance if you ask our model

23:07uh yeah you know what's the what's the better like in france today uh it will give you an answer it will do some tool calls and figure you figure out the answer but in that prompt nothing will signal that oh you know they're asking me a question about uh you know like a root cause analysis on a mainframe uh z system so so it will detect that and it won't activate that part of its training right um but if if the question is hey you know i'm in big trouble you know how do i uh restart you know how do i uh re-initialize this l part on my mainframe then uh that uh that activation layer kicks in and says oh

23:44okay so looks like they're talking about this other thing that i was trained on and it loads those weights and biases um that you create during your training and then those weights and biases uh sway the network just enough um so that it whatever it generates plan or output summarization or what have you is now affected by uh affected by it by that training we found and it's not just us but also this is throughout the academia and now also in the industry the best way to generalize

24:17um a data set um a training data set is really um uh to to go through these trainings so that the model uh can can reason about the native in a model model native model native way um so that's why you know we are uh we've been pursuing this for for some time now yeah uh and then so once the reasoning is done and um and the system has identified uh a problem and and designed a plan of action is that

24:51then passed on to an agent that executes that plan of action yeah um so uh we when we first started we just first we started with uh a textual recipe that we would hand off to the user and say okay so here's what you need to do this is the recipe i need you to follow and maybe there's like eight to twelve steps in there and then we called it a day but of course you quickly realize well okay some first of all some of these actions are um very automatable like once we break it down to a

25:24problem a problem into the little steps it already you know automatically suggest automation so in that regard what we first said was okay so a lot of these are uh first let's verify the problem let's do additional diagnostics essentially a lot of read-only operations uh so we categorized uh we we thought uh we you know we train our model so that uh it knows how to categorize and label these actions as uh well these are diagnostical so you'll never break anything by just doing these

25:55so we started automating those of uh as you can imagine uh some of the actions are on the remediation uh side of things which means um you know oh you need to go go to this configuration and actually change that port number or or go to this machine and allocate more um file systems space so it's actually things that touches the system we still require a human uh human in the loop for them to uh approve the plan uh before any of any any such uh automation automation is actually uh called upon um so the the right now

26:32essentially we are in an um assistive capacity uh which means uh these human beings are now um you know fed these uh recommendations and ai generated plans of action uh and as they do things and as they um as they adhere to the plan or deviate from the plan uh we actually track that life cycle craig so that uh if if any of our um you know recipes are executed and everything is fine that's great that's great feedback for us but we actually learn from our mistakes more uh meaning if let's say for instance

27:05we tell them hey you know you need to do a b and c but then if they do a b and x um then we actually learn that uh once they that once that ticket is closed uh so that we can um we even if that subject meta expert is not really donating and writing a lot of descriptions of things we want to on essentially understand their intent and the the real the real way they they fix the issues uh and we we train our models um as that flag fly will turns as they you know respond to tickets um and and and and and and

27:42leave little little little digital traces of of their actions the human in the loop with a lot of these autonomous systems uh you know there is something called automation bias where people become accustomed to accepting the the machine's recommendation because it's it's worked the last 10 times and so the 11th time you you don't even think about it you just say go ahead i mean is is there a point at which um

28:19in which this will be fully automated and and yeah and the the human uh users will will sort of be monitoring out monitoring out monitoring outcomes uh but but not necessarily uh you know approving every action yeah yeah that's certainly where it's going um there's definitely huge uh economic pressure

Automation and Economics

28:45in uh automation of course um one thing that you said is automation bias is very true um um and to to to address that we we have built various uh quite interesting fingerprinting techniques of incidents and issues essentially uh when something goes wrong uh we keep a very detailed record of uh what the machine signals were what human beings have done um and locality of that issue how that issue transpired like

29:19what was the first domino that felt toppled over what was the second domino so we create these uh causal traces of um incidents and issues in the enterprise so that we can say oh you know what's happening right now is something that just happened last week uh and it seems like every friday you guys are going through this hence it's highly likely that this is it is that issue so it's actually to combat uh that uh and also give confidence people uh that uh you know the this is something they can actually

29:53tackle with automation we uh we have done a lot of work on um fingerprinting uh fingerprinting issues um and uh that uh it's actually in it in this in this sort of problem space it's very important to know whether you're dealing with a brand new problem versus something that happens with some some frequency uh there's that's there's tremendous difference difference between them it's actually being able

30:25to say you know what i don't know i'm throwing the towel in you really need to pay attention to this is is is a great asset you that's not a failure of the product and when it does that we are extremely happy uh that we've been able to do that uh because you know one thing that uh you've noticed probably uh the llms will never tell you oh you know i don't know something right they they're very positive they're uh they're they're uh they they love to generate uh so it's actually to get them to say something like well you know this is really out of my uh bounds is is difficult and that's something

31:00we've been working on uh for for ever since we we uh we started we started this but of course is if you if you are sufficiently um uh if you can sufficiently uh fingerprint these incidents and issues and can sufficiently confident to say hey you know this is actually something we dealt with and here is how we solve these problems uh then there is a lot of uh economy economies there so you can imagine that the pressure in the market is to always towards okay so can we fix these reoccurring

31:33problems which tend to be like 80 85 percent of the problem space can we what can we do to automate these as quickly as possible uh that's essentially what you know people like ryan and me think about all day yeah well one of the things that people have been doing is they build uh this uh kind of a committee of uh llms that that check each other's work uh and then come up with a consensus about you know which what is the optimal uh output or play whether that's a plan and action or or uh yeah are you doing

32:16anything like that yeah we absolutely do uh which is uh interesting uh uh so we um so we have this uh new agent called uh deep root cause analysis agent so this is something that you unleash on big uh problems in a post-mortem sense so uh meaning problem already happened and you're trying to understand what why why that was and you have a lot of time so you have a lot of research research time it's actually it's not a uh it's you don't have to be um you know counting milliseconds necessarily

32:50so when this when this thing is launched uh it uh it uh it quickly formulates an opinion a theory about what it is uh based on all the available data all the snapshot data that was handed to it uh but it has agency so uh in fact you know let me say agentic ai that what what what that really means is hey you know do you have a system that can take initiative uh to dig into a to dig into data and uh come up come up with a come up with an outcome come up come up with an outcome

33:23that is towards that is in the trajectory that you think it should do so for us that's root cause analysis meaning we task our agents to say uh is hey you know here's all the telemetry data that we have here are all the data connectors that you might want to pull in uh you know in an agentic manner if you want and i need you to uh get to the bottom of this and um do a deep root cause analysis on this and when you test this agent it uh it like i said it formulates a um a hypothesis and then this

33:56hypothesis is then passed to other sub-agents in the system uh we have a agent that is expert in log analysis and analysis of log streams and log uh log data in general uh so when when this hypothesis is passed down uh to that agent with the uh with the task because the our planner says okay so i suspect it's gonna be the load balancer again so you know somebody needs to go study the load balancer logs and it hands off that to that agent and that agent uh says okay i guess this is what i do now it

34:33goes to the it goes to the uh log source and then starts analyzing the data uh but we give that agent uh enough enough enough creative space such that if it finds a refuting evidence or supporting or refuting evidence in the data it can uh signal that back to the other agents and say hey you know what you told me was uh was uh to look for this but i didn't find that but i found this other thing which suggests uh your root cause analysis was misinformed meaning i found more evidence for you

35:07perhaps you should reconsider um so we don't have a peer-to-peer-to-peer uh agent uh forum like maybe uh you're imagining but we have a hierarchy of like a principal researcher versus uh individual clerks that are very good at researching different types of um machine machine machine data like logs and metrics and topologies like graphs of things so those are individuals areas of expertise um and they test the hype they know how to test the hypothesis and also create new insights that the

35:42master planner uh can then reason re-reason about and they constantly back and forth you know do this until they converge on on on a response like everybody's happy or exhausted yeah uh the uh uh uh we were talking about how you're containerized and these are built as microservices and you have a list of uh of these services uh from the users the customers point of view do how does that work do they

36:19they choose to use some and not others or is that just uh architected that way uh for because it's easy for you to swap modules out as you uh improve things yeah um so when uh when the platform was first designed it was just like that it was a it had that a la carte feeling of okay so uh here is what i need here is uh let's say uh let's say i'm i run a sre team site reliability engineering team so for my team

36:53i need a workflow like this i'm gonna need this service to be modeled and then i need this workflow to be defined so it it it certainly was like that so that you could uh pick a cherry pick the services that the platform offers you but in the agentic uh world uh you know you know we are quickly migrating uh you know uh everything towards uh agents nowadays uh this actually uh is a lot smoother and transparent from the point of view of the user because they just naturally conversate uh they say

37:26hey you know i'm trying to do this i'm trying to do that and um there is a sufficient intelligence baked into the platform uh baked into the um agentic layer of the platform uh that it knows how to delegate these to the all the other um bots and models and what have you so we have a big um

37:48umbrella uh term you know we call it helix gpt which is essentially all the generative all the reasoning all the analysis tasks that the platform is generating is then routed to uh to uh either a model provider of the user or one of the models we host uh for the user either on their um on premises or on their or in their private clouds versus maybe a cloud offering of choice so this is the flexibility rhino was talking about we actually um you know we are able to deploy uh the components pretty much

38:24everywhere including on on premise um so we have a rather sophisticated routing layer that knows how to you know where it goes uh you know what what workload goes where uh and it does all the metering and all the compliance and and whatnot there's something about the can about what you guys have built with helix that he was saying uh service now isn't able to do that because i that we didn't get to the why i i would say um we have um you know you you can run us in any cloud or hybrid or on-premise

39:04scenario so from that perspective we do have a unique unique stance i suppose um meaning uh a lot of organizations are pretty nervous about um you know sending their most intricate ip and data uh to third party uh ai vendors um because they see how capable this plan and models are they can you know quickly learn things and uh so so there is a lot of interest in being able to run these uh this sort this level of intelligence in an air gap manner if you will um so that so that nothing leaves the uh the or that

39:40organization's boundaries um so for that reason uh we have um you know we built a lot of uh scaffolding around around around these models so that a they they run very efficiently um you know we don't um we actually we we look for parameter efficient models as much as possible meaning um i need to be able to run all this intelligence on highly available not so cutting edge uh hardware um and that does require

40:13uh some work uh some work on like i said being able to uh uh you know the uh train the model so that you don't essentially activate necessarily the entirety of that model for every single token also being able to do um um support uh multiple tenants perhaps on one installation one one single single tenant um so we've done a lot of work to uh to make uh generative and not just possible and exciting but

40:43also economic to run um so from that perspective i think we do have uh like i said i don't do a lot of uh competitive analysis as part of my job um but from that perspective we know i know um we have pretty unique um features um and advantages uh yeah and and and if somebody wants to to run this on premise or a part if it's a hybrid solution you know on premise and in the cloud how do you deliver the

41:16model or the modules to them is that just they download it and install it on their side uh pretty much yeah just so uh so you uh you know we talked about containers and containerization uh so this this happens to be um yet another set of containers uh you you deploy as part of the uh the product uh the only uh perhaps maybe slight difference is uh the the machine that that's going to run this container

41:47will now need access to a uh a decent gpu um and uh we support uh pretty old uh hardware all the way from uh nvidia's uh l4s uh to a a100s and nowadays rtx 6000 are are pretty good so we um so we just require them to uh procure or lease or you know find find a gpu uh and then um we give them the software architect

42:19that basically is part of our platform really uh and that thing turns that turns our model into an inference server that all the other things in the in the system can then start utilizing um the uh and also the training we do we have we have this optional uh tenant specific training pipeline also uh which what what that is is you know they could just go to our administrative ui and tell us where their sources are where their ticket sources are or their knowledge articles run books and what have you

42:52and what we do is uh we harvest all that and then we analyze the data to see um what can be learned from them um so that's also pretty i would say interesting uh because that allows us to uh uh to keep up with the organization because no it organization is constant right so they deploy something maybe their network is all cisco hardware uh but all of a sudden somebody introduces new juniper hardware somewhere so they they shift right and as they adopt new technologies and new um you know new components all the failure

43:27modes change um they they bring their own problems they bring their own compatibility issues and whatnot so there's always movement in uh that we have to deal with uh so for that reason we tap into these uh sources and periodically pull the data out to see um what needs to be uh generalized and what's what's learnable from them um so anyway so that's also another uh component that goes with this um generative AI uh inference server uh and once you once you do that you have uh you uh you periodically let's

44:00let's say every 24 hours your model revs up so your um it it might decide to learn new things and and update update itself uh that way you're uh it's it's never behind um so yeah yeah um the um

Performance Metrics

44:16um do you have uh metrics of uh you know that that demonstrate how uh this helix performs uh in a large organization i mean these are big decisions by uh by organizations i imagine that much of your customer base has been with you for a very long time but to get somebody uh fortune 2000 to switch

44:46i mean they've already got a solution um to switch to uh helix uh is a is i would imagine a long sales cycle so uh do you what what do you present to them as uh evidence that that you guys save them time and money yeah yeah yeah so uh yes we always uh um you know we have to uh refer to um uh some some

45:18measured numbers some customers but what happens craig at the end of today they it always comes down to a bake-off of sorts so they um uh just like you said these these organizations um don't buy these uh tools just to use for a year right they typically plan for three to five years out uh and it's often the case that they ask us okay so uh you know what can you do now but we also need to understand where you are going you know to see if if our goals align uh so just like you said these are actually very

45:50complex long uh sales sales cycles um what we uh what we tell them is uh so we do have a portfolio of um uh savings that measured measured savings that we share with them so this typically based on the use case this typically anywhere from uh 25 to 40 sometimes 50 percent so you know we we give them these numbers uh but of course at the end of the end we have a lot of benchmarking information uh both uh

46:23synthetic and open source uh and third party you know we also uh we have done uh some we have competed in some competitions just to show people you know what what this what these models can can can do uh in isolation uh so we have a lot of uh proof points uh proof data but at the end of the day um like i said it's often what happens is you get deployed in a um pre-production environment or some some qa environment that they're they're choosing sometimes together with your competition for with other tools

46:56and then they just measure they just say okay so uh you know how given an operator and this tool uh and the day maybe you know they sometimes uh you know they inject the problem and then they see okay did you find it did the operator did the operator gain any insights from you or or so on and so forth so it's actually uh uh this this sort of uh software i've never seen you know be being sold unless you actually do a proof of concept uh so that you call that you collect some uh you collect some data

47:26yourself and and on where you're going i mean what do you what what is on the road map um so uh you

Future of Operations

47:37i mean uh you you see how software engineering got disrupted right software engineering uh i would say within the last six to eight months um is now different fundamentally different right we had tectonic shifts in software engineering uh and uh we we believe a similar disruption will also happen in operations in ops ops um and the reason is you know software engineering is a very verifiable problem right so you can generate a piece of code and then compile that code and say okay does it compile

48:13or compiles does it run you can run some tests against it to see so it's like a all you can imagine you can create uh training flywheels for software pretty easily and that that's why that was the first thing that got heavily heavily automated now uh we believe ops are also close to this i mean it's the uh verify verification cycle of software is maybe measured in minutes for ops that's typically hours um in maybe perhaps

48:44days but at the end of the day it is still a verifiable problem because these are digital systems in at the end of the day and it system is uh it might look complex i mean it's a complex system but it's not a chaotic system right so at the end of the day these are still computers running software and you know it is at the end it's a deterministic it's a deterministic system but the question is can you verify this in uh quickly enough to be able to train your model and that's exactly what we are doing

49:17right now great so right now uh most of our research effort is going through uh building uh gyms or labs if you will uh where the agents can go and get trained on um so uh large language models were uh a function of text data that was available on the internet or offline right so they're they gone through piles of piles of text data uh but we are probably at the limit of that data so nowadays in the future uh what's going

49:47to happen is uh people will start exposing these models to the real world more and more so that they can create they can gain that first person first person view of things because uh right now ai if you think about ai is only learning through someone else's description right text means you know someone wrote something and that's their interpretation of reality right so so now everybody's trying to cut that off and expose the ai directly to the reality of whatever domain they are in uh so that they can um

50:20you know they can uh do things and learn from that and in our space that means creation of these what we call gyms so gyms are essentially imagine a data center um that just runs whole bunch of different applications but maybe nothing is sensitive or created for for the for the purpose of training in the sense that i can let my ai go there and break things i can say okay so create like a chaos monkey agent that just goes and randomly breaks things and say ah you know i took i i i messed up your

50:52application which then my other ai comes in and says okay uh so let me first verify what's going on let me try to fix this and it tries different things and i log all of this uh i i keep a record of how how things are progressing and uh whether the remediation works is it expensive is there a better one so i can do all these in a 24 7 lights out manner in a data center so that i can create a bespoke agent that knows all the intricacies of whatever application it's responsible for

51:26um so that's i think that's very that's very obviously that's where we are trying to go i think that's where the industry is also going towards so you'll see more and more you'll hear about these you know world models and uh whatnot that's also sort of uh philosophically aligned with uh what what i just said um but uh what i can i mean what i can confidently tell you is these things are a function of their training data uh and good training data via we exhaust textual space uh so most of the best data will probably come from um environments that are created so that these these

52:06ai agents can just go and experience themselves and learn from the learn from those mistakes yeah yeah and that's fascinating i mean i've talked to uh a number of people about world models and and direct uh

52:23learning from uh reality as opposed from text uh so it's exciting to hear that that research is going on are are you developing your own models in in uh for that or or do you use uh uh uh uh you know models from one of the uh foundation labs and then find it yes so we have always used a foundational model um both in our research and in our shipping functions um and um

53:02because they are just marvelous and good enough on their own they uh and also you know training a foundation model is pretty uh pretty expensive and slow operation uh so we always started from uh foundation models uh but we have a uh pretty uh um advanced workflow now i would say that that allows us to uh try out new ones all the time so i i have a radar that sort of you know we constantly try you know different um you know new um new foundation models from different vendors uh benchmark them train them

53:37uh uh and and and and and to to evaluate them these days uh the uh the uh the uh we we are really um found of the quen family from uh alibaba uh there they they do great great uh uh frontier work um and um the uh also gamma 4 from google is is other uh so we have two editions of the same model um so we have a um a coin based model uh and then a google google based one um and we offer them both

54:13uh and and the foundation model is very uh the way our uh we are architected is swappable uh so so as these things um you know develop and advance um we are not really married to to them very very very tightly so we can easily uh swap them out uh in and out um it's a it's an empiric so it's interesting craig that it's it's a very empirical uh space uh so you can't just look at the architecture you can't read the paper and say oh that's a better model uh you really at the end of the day you have

54:48to sit down and uh and test these uh try these uh there's all kinds of reward hacking models do you you know when you're trying to test i mean some sometimes they detect that they're you know oh it looks like i'm being tested right now so uh it it's it's it takes a different rigor uh to formulate an opinion about a foundational model uh but that's we are all learning that you know i think everybody in the industry is learning some techniques around uh being able to evaluate this and essentially right size them uh because you'll you'll see that one model comes in a whole bunch of different um sizes

55:22and resolutions and quantizations and what have you um so uh so you you know we did have to tool ourselves and so that we can make these judgments uh you know data driven data driven way but like i said we we are we offer two foundational models today but this is uh the future of it service uh management and and it sounds like it's going to get increasingly automated is what's that going to do to uh to uh to uh to it staff does it make them change what they need the skills that they need

56:02uh or is it uh is are they going to become agentic managers or managers of agent systems as as opposed to having their fingers in the code so the good news is i think that what's going to happen is the quality will go up um because uh what we what these systems these agents really tackle target is repetitious things that shouldn't shouldn't even have happened for the for first place you know

56:35there's a lot of um uh you know we expand a lot of calories uh just to fix the same problem again and again uh or things that could have been prevented to begin with and uh when we uh when we work with our customers or our own it personnel you know who are you can imagine they're quickly adapting all these technologies also they actually uh so right now my reading is they're quite happy to to have all this assistance they say well you know this saves me a lot of time i can go uh you know go home at a

57:08a predictive uh predictable time now i'm not swamped of all this garbage fire firefights i can now concentrate on actually making the organization better or making my such and such projects better uh so so so this um you know this sort of technology buys you additional digital capacity as you know ryan calls it uh to then invest in other other new things it's actually i think the smart growing organizations will take this digital capacity and then redeploy it in uh solving the customers or end users uh

57:45problems better or offer them you know more services and more um um more more digital goods uh and uh what the what does this what means for this individual is people who have um agency is now able to move a lot faster uh so we see this uh uh you know the certain cohort in every organization sort of really shining uh you know with with the help of with the help of these uh

58:17technologies uh so i think some people uh will move up in the chain some people will be disillusioned and uh will say oh you know what i'm done i'm retiring uh it's certainly a big economic labor uh labor shift you know it's happening in front of our eyes uh but my like i said i don't uh my personal opinion of this is this is great for the society i think everything will get higher quality and will become cheaper because of this uh and for the people who are involved i think they'll they'll just uh

58:50they'll just tackle uh higher order higher value adding uh adding problems uh as these technologies get bigger and become more ubiquitous uh as they will yeah um okay

More from Eye on AI

Why People Are Paying 10x More for AI | Sid Sheth, d-Matrix

Aug 6, 202650 min

Real AI Transformation Costs HALF of Everyone's Salary for 2 Years | Chris Blackburn, Liatrio

Jul 30, 20261h 5m

"According to NASA's Definition of Life, I'm Not Alive" - Why Nobody Can Define Life | Dr. Kate Adamala

Jul 29, 202646 min

Video Is About to Stop Being One-Way (and That Changes Everything) | Victor Riparbelli, Synthesia

Jul 28, 202639 min

Video Is About to Stop Being One-Way (and That Changes Everything) | Victor Riparbelli, Synthesia

Jul 28, 202639 min