Watch the ComplianceCow CCM 3.0 Webinar
This 30-minute webinar is based on a session ComplianceCow was invited by ServiceNow to present in the AI Innovation Zone at Knowledge26.
The presentation shows GRC teams how to go from a plain-language prompt to a production-ready compliance control, in minutes. It shows how Generative AI can help compliance teams define, generate, validate, execute, and document automated controls without long implementation cycles or heavy engineering effort.
Watch the replay to see how ComplianceCow turns plain-language control requirements into production-ready automated controls.
From Months to Minutes
Traditional CCM implementations, whether RPA-based screen scraping or API-driven integrations, often require significant engineering effort and long deployment timelines. GRC teams struggle to keep pace with modern DevOps cycles, spending more time chasing manual evidence than managing risk.
CCM 3.0 addresses this directly.
Using Generative AI, compliance teams can describe a control in plain language and have structured, production-ready control code generated, validated, and deployed in minutes. Outputs are deterministic, meaning the same input produces the same output every run, and fully auditable, with a complete trace from prompt to evidence.
What You'll Learn in this Generative AI CCM 3.0 Webinar
- Why CCM 3.0 represents a major step forward from prior continuous controls monitoring approaches.
- How Generative AI can produce deterministic, auditable outputs for compliance automation.
- How to move from a natural language control description to a running compliance control inside GRC platforms, like ServiceNow IRM.
- How automated evidence can be generated and attached directly to assessment controls.
- How ComplianceCow helps GRC teams turn plain-language control requirements into production-ready automated controls.
Speakers
Raj Krishnamurthy — CEO of ComplianceCow
Megha Shah — Head of Customer Success
Webinar Transcript — From Prompt to Production: Generative AI CCM 3.0 Controls
Read full transcript ▼
Shradha Krishnamurthy: Thank you all so much for coming to join us today. This is an iteration of a talk that was given at ServiceNow asked ComplianceCow to make at Knowledge26 as part of the featured AI Innovation Zone, speaker series. So Raj and Megha were down there giving this talk. We've modified it for a webinar audience. We're super excited to get this started. I'm Shradha Krishnamurthy, I am a Product Marketing and Operations Specialist here at Compliance Cow, and for this talk we'll be featuring speakers, Raj Krishnamurthy, who is Engineer Zero and CEO of ComplianceCow, and Megha Shah, who is Head of Customer Success here at ComplianceCow. Megha manages all our forward, deployed engineers. So, with that, I'm really excited to turn over the stage to our hosts, our speakers. Yeah, super excited for this webinar. Let's get it started. Raj, I'll pass the mic over to you.
Raj Krishnamurthy: Shradha, thank you very much. What I want to do is that… today's agenda, we are going to talk about CCM, Continuous Controls Monitoring 3.0. Why Continuous Controls Monitoring 3.0? What is this terminology? There's something that I think it doesn't exist, so we have somewhat made up, so we're going to talk about what that is. But what we want to also do is to dive into a very specific example that we can walk you through… of what an agentic continuous controls monitoring looks like, which is what we are calling CCM3.0. So, before we dive into the details, what I wanted to do is to first start with, something that you might have heard of. Some of you might have heard of the Mythos project which is on pause right now, but eventually it'll come back again, online again. And you might have heard of this project called Project Glasswing, right, which is a… which is a project that was started with the private, Mythos model, and especially with a consortium of industry leaders, like Palo Alto Networks, Microsoft, Apple, so on and so forth. Few weeks earlier than that, some of you might have noticed a talk by Nicholas Carlini. He gave this talk at Unprompted 2026, where he was basically using a cloud model, to be able to automatically determine vulnerabilities. And the way that he did this, rather, was super simplistic, right? He basically wrote an iterator, which iterated through every single file system on the Linux system, Linux kernel. And it identified some of the vulnerabilities there. And some of these vulnerabilities have been latent for the last 27 years. And some of these are not direct logic, I mean, they were essentially trying to make a very complex multi-step process, so the models have become sort of really sophisticated in inference and understanding. And it is very fascinating, as you sort of look at what he did, that some of these vulnerabilities that he determined that has been there for a long time, and he was able to do this very, very quickly, right? So, what does that mean? Why is this discussion of Fable 5 and Mythos in a discussion around continuous controls monitoring. I think the question is, what does that do We all understand the sort of the pressure that it does for the security teams, but I think on the other side, what does it do for the governance, risk, and compliance teams? What does it mean for cybersecurity risk? What does it mean for cybersecurity compliance? And how does that sort of modify in terms of how we think about these things? And that's why I started this conversation with…with, with my task. To take a step back, let's just make sure that we are all level set, right? We see GRC as a tale of two worlds. There is the world of what we call the first line of defense, your engineering teams, your IT teams, and your security teams, and security operations teams, SecOps teams, and security engineering, so on and so forth. They are operating on a much more rapid clip at the speed of business, right? This was true in the cloud era with a lot of this in terms of the acceleration of applications, but it is… it has gone exponential, right, in the GenAI world, right? Which is… in which 80, 90% of the code is auto-written. And the verification part is predominantly done, at this point in time, done by humans, and eventually that may also be taken by agents. So the rate at which we deploy applications has fundamentally changed. So the cadence at which you are talking about these applications have completely changed. While, on the other hand, if you think about what we are doing with risk and compliance and audits, we are still sort of stuck with some of the cadence, right, that we are used to. So the idea of identifying risks, right, or evaluating them for compliance, the cadence mismatch is extremely apparent. And these two… so that's what we mean by the tale of two worlds, right? I mean, the world of the first line of defense operating at a completely different place than what we call the second and the third line of defense. And that, basically, I think this is not a new concept, right? I think we have really been realizing this chasm for quite some time, and I think if you really look at the last decade. A big part of that was around, you know, how do we sort of better automate? And by automation, what it basically meant was how do you take better screenshots, right? Can we take screenshots fast enough? And that's what we call the CCM0.2, right? So if you were having this conversation 10 years ago, if I say continuous controls monitoring, most of us will be talking about, you know, RPAs, right? The robotic process automation, and how we can take screenshots faster. The idea was to mimic the exact human process that existed. To be able to perform this at a much faster clip. Obviously, that doesn't scale, because as the APIs became…sort of the standard, right, consuming for a protocol, and the software development kits, sort of… and the legacy applications start migrating, and you've started getting into the 12-factor apps and the new modern-age applications, almost every single application today has become API-centric, right? Now, the question then becomes, and I think that's something that we have been dealing in the last, I would say, few years, I would actually, I would not necessarily put it at 2020, it has started even beyond that, but I think from a maturity perspective, I think you have started seeing this in the last few years. Which is that, how do I use the APIs as long as I can demonstrate, sort of, the chain of custody, right? As long as I can demonstrate, the fidelity of the transactions that we get. How do we use the APIs to go solve for this continuous controls monitoring problem? Right? So no longer just relying on the screenshots, maybe we'll still do some screenshots from a sampling perspective, but how do we use this, and how do we use APIs to integrate and collect data? And I think…that has been working fine, I think, meaning that… I think most of us are still maturing into that phase, but I think the agentic era has completely changed this, right? The rate at which agents operate has completely changed. And there are some challenges with CCM 2.0, and I'll talk about them as well, but the idea is that how do we now keep pace with this exponential rate of change that is happening in how applications are built, deployed, and how do we want to assess risk and compliance for those applications, right? And that's what we call CCM3.0. So back to the main topic, when we say CCM 3.0, what we are basically saying is that, how do we do continuous controls monitoring? How do we re-envision, sort of, what governance, risk, and compliance is for this agentic era? And that's what we mean by CCM 3.0. Now, what I wanted to say is that, like, the CCM API-driven applications were effective in the sense that you no longer were doing sampling anymore, you were looking at the total subset as much as possible, right? But there are some challenges, right, trying to get these APIs to work. And until very recently, most of us would have gone through this. It takes… because these are rule-based automations. It takes very long to be able to develop and deploy these APIs, right? A lot… I should say relatively longer. I mean, they're not… they don't necessarily… I mean, you can write some of them very, very quickly, but it takes, definitely on a relative basis, much longer. You need to have… to be able to go develop and deploy this. The second thing is that most GRC teams were dependent…on engineering teams, and this is where the entire notion of, for example, GRC engineering, like I said, discipline came about. I want to do a big shout-out to folks like, so Mosi Platt, Ayu Fandi, a whole bunch of others, right, in the space, who had sort of pushed this idea of GRC engineering as a discipline, and that is where I think this idea of GRC, sort of trying to engineer this, right, from an automation perspective, came in. But the problem then, again, is that in many cases, we were dependent on engineering teams to go solve this problem, right? So there was a skill gap. Or a skill retooling issue there as well. The third thing is that, I think the… the…screenshots continued to exist in some cases, and I think there was a sort of a mix of, you know, how much of APIs and how much of screenshots, so there is… there was definitely sort of a manual fatigue that also happened there. But more importantly, I think the controls drift is a much… became a common phenomenon in this 2.0. So how do we make sure that we keep up to, you know, the rate of change of all the APIs that are happening? How do we demonstrate the fidelity of the transactions as it moves from the source systems all the way through compliance? Into the audit systems, right? And those were some real challenges. Now, the main thing that changed, in my opinion, I think particularly large language models, as you guys know, and I would apply the word thinking models here as well, because the language models have become a lot more sophisticated, especially with the mixture of experts models that are being developed. Right, so what has happened is that we are no longer dependent on static rules and writing static code. I think we are now able to infer and discern semantic inference right at runtime. So I can understand what the user's intent is, I can understand what the data says, and I can start, sort of, dynamically building these things in the new world of LLM. So the question then, that leads us very nicely into CCM3.0, but that is also not without challenges, right? I mean, the new challenges that we face today, where we have moved from the speed at which, the breakneck speed at which we can write and consume and build around these APIs, that is good. But there are some other challenges that we'll have to deal with, right? One is that agents are inherently probabilistic. I'm sure you guys all know that. The problem fundamentally also becomes that because agents are chained, you are compounding these probabilities, it makes it worse. So, 90% of 90% is 81%, right, as you all know. So, as you chain these events, right, they become more and more non-deterministic. There are ways to solve this, which is what we are also going to talk about, but…they're basically fundamentally sort of trying to solve this purely looking at this from an agent, right? If you're using your cloud desktop, CloudCover, Cloud Core whatever that is, right? OpenAI, and they become inherently non-deterministic, and you'll have to, you'll have to figure it out. Right, so idempotency, as it typically existed — idempotency, just to level set. For the same set of inputs, unless the underlying state conditions don't change, you're expecting the same set of outputs, right? And that you cannot guarantee in an agentic world, in a large language model world. The 2 is the chain of custody. Because most of us… you can actually go try it out today, right? You can go try and write, describe something, and it'll maybe put out a code, a sense of code, in order to be able to go connected, but the problem is that it becomes very difficult for you to demonstrate the… the… make it auditable and to be able to make it, demonstrate the chain of custody, right? In terms of how…the auto-code, auto scripts that you develop, right, through agents, how are you going to demonstrate that they are fully auditable? And that's a big challenge that we'll have to sort of deal with today. The third thing I would say is that, and I'm grouping a whole bunch of things into operations, and maybe it doesn't do justice to put all of them into one bucket, but for the purpose of the brevity and the simplicity of this conversation. It is not enough if you simply write scripts. The question then becomes, and that is fantastic from a demo perspective, right? We can quickly demo something. The question is that, how do you deploy them at scale? And the argument that I would make is very similar to writing a Docker file versus running a cloud-based application. Very different set of ideas, right, because at scale. You get into a whole bunch of distribution problems, and systems engineering problems that you have to deal with. So, it is one thing to go, you know, write an API to go create a pull request and evaluate a pull request, it's a completely another thing for us to be able to go scan this across hundreds and thousands of Git repos and the pull requests in order to be able to say what is compliant, what is not compliant, what is conforming, what is not confirming, right? So the question then becomes, how do you deploy them? How do you execute them at scale? How do you maintain these brittle scripts? How do you create observability? How do you continuously monitor them? These are problems that are fundamentally becoming extremely difficult in the world of…you know, writing agentic code, and you'll have to solve for them as well. And finally, I think often overlooked is what I would call the token sprawl. I think it's an understatement. But fundamentally, if you really think about the cost that you're incurring, right, these are staggering costs. The Frontier models cost, for example, somewhere around $50 in, $10 in input tokens for a million tokens, and $50 in output tokens, and if you sort of do a simplistic math, right, you're looking at hundreds of thousands of dollars, right, across a couple of users within an organization, and you have to be prepared. And if you… as long… if you don't control them it may… the cost skyrocket. We have seen through the iterations of this before, with the cloud sprawl and all of that, and this is no different, but you have to be prepared with this, and you need to have the right controls in place. But more importantly, it's beyond just the dollars. You have to also account for the fact that you're going to run into rate limits, right? At a certain point in time, cloud is going to turn you off. And you have to understand how this is going to impact productivity. The last thing I would say is that, in my opinion, sort of going out of scope on this, it is also not clear how much of these costs are subsidized. And what would be the long-term, you know, especially if you look at these model companies right now, and sort of their operating margins, we also don't know where, eventually the steady state is going to look like, right? Would it continue to go down, or would there be other costs that'll get added that we are not seeing today? We don't know a lot of these answers, but what we do know is that even with the answers that we know, token sprawl is a very real thing, the productivity impact due to rate limits are very real, right? So, this leads me back into the discussion around CCM3.0. Where I think in order to solve this problem, what we… what… I mean, ComplianceCow isn't the… I mean, we are primarily built and we, from a… we started as an API runtime for GRC, and we morphed ourselves into layering, agentic, you know, agents on top, yo be able to go solve this problem. And the point of view that we took is slightly different, right? What we said is that the large language models inherently solve what I call the long-tail last-mail problem. It allows us to be able to describe very efficiently in natural language to be able to automate, but you need the primitives, the basic building blocks. Like, AWS offers you EC2 and S3 buckets, and your network policies, and a lot of these different artifacts and primitives that you can bring together to construct your application. We fundamentally see there are building primitives that are required for governance, risk, and compliance as well. And using the large language models, it allows you to design, but to be able to deploy deterministically at run scale. So, going back, what we are basically saying is that it allows us to be able to solve some of these problems, where you can build them, right, and allow the user to be in the loop, and validate them, but you'll be able to deploy them to produce deterministic output, right? And you can sort of move away from this idea of non-deterministic. So, to give you an example, if I have to go evaluate a particular rule to collect a set of deployments from my CI/CD pipeline, or from a given Kubernetes cluster. Then it's easy for me to model it as long as I'm not hard-coding any of my application credentials or anything like that, and now I can go deploy what I modeled, the template, and go evaluate this across thousands of clusters, right, without necessarily being… and being fully deterministic. Two, it creates a very clear chain of custody, because you're essentially curated what you're collecting the data, and you can demonstrate the entire auditability. And three is that it becomes… you handle all the associated problems with it, right? You know, how do you deploy at scale, which is the fundamental issue we talked about. How do you maintain, how do you create, execute, so on and so forth. And finally, it becomes… it actually lends itself to the token sprawl, because you are only limiting this to the design of a particular artifact or a design, and you're not necessarily trying to consume tokens through all of the data, right? Which is practically and economically not possible at all. So, what, I mean, whether you look at compliance code, or whether you look at any of the products, what we would strongly sort of recommend is that please ensure that you have the fundamental building primitives, right, in order to be able to go solve the…the automation of evidence collection, controls testing, potentially even remediation, and make sure that you have the right building primitives to be able to go solve those problems, right? Before anybody says Agentic, I think the question you should ask is that what do these agents rely on? What sort of runtime do you provide? And that should be the standard questions that you want to ask. And we have actually taken some of this, and we have open-sourced this as well. I see some in the audience, you might already know this. You can go to openssecuritycompliance.org. You can Docker Compose up, you can bring your own large language models into it, right? In fact, you can even do DeepSeek, not that I suggest recommend DeepSeek, but if you want sort of a cheaper way to try this out, you can bring your own models, and you can start building some of this within open source as well, which I… which Megha is going to give you a demonstration of, right? So, we believe that the future of risk and compliance, as it rightly says here, is not about writing rules, but it is about describing intent. And I think we've made a lot of progress. And I would say that we are a very step closer to what I would call a declarative GRC. Very similar to what we see in cloud declarative systems and Kubernetes declarative systems, where we can talk about the to-be state, and we can reconcile that to the asset state. With that, I'm going to hand it over to Mega for walking us through a demo. Megha, please take it away. I'm gonna stop sharing, Megha, and if you want to start sharing…
Megha Shah: Okay.
Megha Shah: All right, thank you, everyone. Thank you, Raj.
Megha Shah: For setting the context, so we'll… today we'll take a use case where we can, we'll basically try to answer whether we can audit a zero-day vulnerability, and the steps that we'll take for today's demo are, as a user, we'll prompt the rules agent to fetch the list of vulnerabilities, specifically Kubernetes vulnerabilities. Next, we prompt the ComplianceCow rules engine to create a compensating control to identify the clusters that have these vulnerabilities, and also give us the steps to remediate those vulnerabilities. So, with that let's get our hands dirty and dive into the demo. So now, this is the AI Studio under Compliance Cow. An open source version is also published in the link that Raj shared earlier. You can also execute this from Claude, Goose, or any other MCP server, right? So we'll start with listing the latest Kubernetes. CVEs. Now, you have, the AI gives me the list of different CVEs. Let's zoom into the first one. So, what it says is ingress, this one is related to an ingress NGINX validation admission controller webhook, and what it means is any person, unauthorized attacker, can get an arbitrary access to my Kubernetes cluster, it can exploit the cluster, it can read the secrets disclosed within the Kubernetes cluster, it can fully take over my cluster. So that sounds scary, and you don't want this to be sitting unnoticed, right? Now, the… These are typically found… this… this is typically found in V1.11.4 and 1.12.0, and the remediations are you either upgrade… so permanent fix is to upgrade to a particular version of Ingress NGINX, and then as a temporary fix or a workaround, you can even uninstall, or you can disable, right? The web… the webhook. So, with that, you have a list of vulnerabilities. We zoomed into the first one. Next, let's take this first one and ask the Compliance Cow Rule Studio to create a rule. Now, rule in Compliance Cow is a series of tasks that are piped together to create a systems workflow. Okay, so we ask it to create a rule which assesses this vulnerability on my in-scope Kubernetes clusters. So the first thing that the rules engine does, it first does a lookup. If it already has a rule. In the catalog, fine. If not, it will start building it from scratch. Like, in this case, it gave me a game plan of these 6 tasks. Now, what are these tasks? These are pipelined together. The first two will give me the list of deployments, daemon sets that are deployed within the Kubernetes clusters. The second will give me the actual configuration that is sitting in my Kubernetes cluster for the validating webhook, and the third and the fourth will flatten the data.
Megha Shah: The fifth is going to do a join on both the deployment as well as the validating webhook, and finally, the sixth one will tell me, it will evaluate the CVE posture on all my resources, and it will tell me whether a given namespace within the cluster is vulnerable to this CVE or not, right? Now… It will create… what happens in the backend is that the rules engine creates these tasks in the backend. Now, these are modular tasks created by Compliance Cow. You can create your own task and publish them. In the backend, these are Python or Go code blocks that are weaved together to write your automation. We have these tasks. You also have the inputs to these tasks, and you have the outputs, which is nothing but the compliance report, okay? Now, that we… what we just did is we listed a vulnerability. We automated the process of identifying the clusters which has this vulnerability. Next, we'll unit test this rule. Okay, so I asked… I prompt my AI studio to execute these rules. Okay, so it will ask me either I can provide the credential here, or we can… pick the credential from the available application catalog. Now, what it does is that you can monitor the execution of these tasks, and at this point of time, if one of the tasks fails, you can even troubleshoot prompt the AI to fix it, okay? Now, say you don't agree with the logic or the task that it produced, you can tweak the logic, you can prompt it to tweak the logic, and make changes at this point and time. So what we are right now doing is we are designing the rule using the probabilistic model, right? And we are unit testing it. Right, now let's… these are the unit test results, which are executed on a handful clusters or namespaces. Let's zoom into the results. The first one is about a namespace. It says it's non-compliant. It gives me a reason, because, see, the webhook is active, but the version is not up to mark. So that's why this is inactive, and this is a very critical risk. The second wasn't compliant, because though the webhook is active, it has a good version. It is up to mark. So, yeah, we are not at a risk. Let's see the third one. Though it says compliant, there is a flag, there is a medium risk. So, here, the webhook is not, enabled. It's inactive. But still this needs to be patched to the latest version, right? So you have a medium risk, and so the last one is critical, right? So that's how you have the recommendation, and you have the rule execution results. Now, vibe coding is good so far when you have one cluster, 2 cluster, 10 clusters. Now, imagine the number of cluster goes to 1,000, multiply by them by the number of namespaces, multiply them by the number of deployments, right? Or the resources in your cluster. Your infrastructure is going to fight back. You will run into all sort of throttling issues, rate limit issues, like Raj was mentioning, in case of Git, right? Same way, even for Kubernetes, you will face all of these issues. So how… what do we do next? So we use… we leverage the probabilistic model to design and test the rule. We then deploy it, and then we run it in production without AI, so we pipe these tasks, we let the orchestrator do its job, and every single time, you'll get a deterministic output on a real time. So that's how you automate, the controls using CCM3. That's what we coined the term CCM3.0. Okay? Now, let's do a little bit of comparison here. So, CCM 2.0, you're already doing automation, very good, but to automate a control, you need to file a ticket, wait, wait for weeks, engineering team is already burdened with a lot of work, and this is an additional perk for them. In 3.0, you can describe your control narrative in a plain language. And you can get done in minutes to few hours. Now, who can build it? Only developers who know to code Python, Go, Java, Rust, right? But here, a GRC analyst is empowered to develop these controls themselves. Now, what happens when your requirement changes? Okay, you've got your engineering team to work, you have… you got a good release, 2 weeks, 3 weeks, you got the control deployed, and the organization decides to move from Jira to linear. You're back to square zero. You now, again, need to file a ticket re-scope, rebuild their weeks of work and weeks of effort lost. But here, with CCM 3.0, you can simply tweak, prompt it to point from Jira to linear, and your rule gets done before lunch. Now, control quality check. Again, you need to describe your requirements, your… what scope to the engineering team, and a lot of communication needs to happen, and it depends on the understanding between the GRC as well as the engineering team. But here, you get a better control, because you own the requirement, you own your control output. And from coverage and scope perspective, also, it again depends on the bandwidth the engineering team has, but here you can automate every control, every domain, every possible framework which are in… for your… which are in scope for you get with that. So the key takeaways of CCM 3.0 is we are expediting the process without sacrificing the quality, so we are leveraging the best of both the worlds. We are leveraging the GenAI speed to design the control, and then we are deploying the con… the rule runtime at scale, right? To get deterministic output every single time. Now, RPA was simple but brittle. API is slow, but then GenAI, it delivers results, fast at a DevOps speed. And then, yeah, vibe coding is… it empowers the GRC team to have a better control output and evidence, which are audit ready. With that, I hand it over to Raj.
Raj Krishnamurthy: We are 3 minutes over, so thanks for staying on. I'm not sure if we have time for questions, but please feel free to hit us. I don't know, maybe if you can put up the slide, Shradha, you can hit us on LinkedIn, you can hit us on email, happy to answer any questions, thank you.
Shradha Krishnamurthy: Thank you so much. Yeah, if you're able to scan that code, I'll also be dropping their LinkedIn here in the chat. You're also… there will be a email that goes around with the recording of the session, some follow-up info, and I'll be sure to include how you can connect with our speakers there. Thank you so much. Oh, David has a question: “What is a deployment model, and how does one try this?”
Raj Krishnamurthy: I think the deployment model, first, I would say that if you are… if you, if you're comfortable, you should go try out the open source first, openssecuritycompliance.org. And you can deploy this on your own, and try this out. ComplianceCow on the commercial side, or any CCM 3.0 deployment model, in my opinion, should be fairly straightforward, because your interfaces are completely changed, and you should be able to Do this natively with Claude Code, or Cowork, or anything that you're already working on. So from an end user's perspective, your expectation on CCM 3.0 should be very seamless, but from a back-end deployment perspective. You can go try out what we built in open source, you can also… We can also stand up a domain for you fairly quickly. Okay, we have no other questions, thank you all.