[00:00:00] Good afternoon, everyone.
[00:00:03] Karasoft would like to welcome you to the Agentic Data Engineering
[00:00:06] When AI Agents Build Trusted Production Pipelines Karasoft Webinar.
[00:00:11] Karasoft Technology is a trusted government IT solutions provider
[00:00:14] delivering software and support solutions to federal, state, and local agencies.
[00:00:19] Karasoft maintains dedicated teams to support sales and marketings
[00:00:22] for all its vendors, including Clearfracture.
[00:00:25] At this time, I would like to introduce our speaker for today,
[00:00:27] Brian Frucci, Chief Technology Officer at Clearfracture.
[00:00:31] And with that team, the floor is all yours.
[00:00:34] Thank you very much, Kylie.
[00:00:35] So, like, I'm Brian Frucci, I'm the CTO here at Clearfracture,
[00:00:38] and I really do appreciate everyone making a little bit of time today
[00:00:41] for us to introduce our Agentic Data Manager to the community.
[00:00:48] A little bit about my own background.
[00:00:49] I've been the CTO of Clearfracture for two years,
[00:00:51] but I've been doing big data exploitation for over two decades.
[00:00:57] Just prior to Clearfracture, I was a tech lead for the Chief Digital and AI Office for Europe,
[00:01:03] bringing AI to the fight in Eastern Europe.
[00:01:05] Before that, I was the CTO of Big Bear AI, took them public in 2020.
[00:01:10] Claim to fame there was we built the command and control AI that was operating the Navy's
[00:01:16] drone task force, 59, the fleet that is operating in the Persian Gulf right now.
[00:01:22] Before that, I did quite a bit of work around the IC, specifically with a company called Indeka Technologies,
[00:01:30] which was Elasticsearch.
[00:01:31] Before there was Elasticsearch, and then I was in the Army as a signal officer even before that.
[00:01:35] So quite a bit of time in the national security mission space dealing with national security mission problems,
[00:01:41] and the amount of labor that we put towards those big data efforts is why we founded and are building out Belvedere,
[00:01:52] this Agentic Data Manager.
[00:01:54] So I'm going to walk you through that.
[00:01:55] We're going to do a couple of slides and then a cooking show demo where you can see the application live.
[00:02:00] And then, of course, I'd love to field questions as we go.
[00:02:03] So with that, I'm going to dig into the slides here.
[00:02:07] So the reason that we are talking about Belvedere is that data is dumb, right?
[00:02:15] Data doesn't operationalize itself.
[00:02:18] It doesn't tell you about itself.
[00:02:20] It doesn't exploit itself.
[00:02:24] And we have got to do that with tools, but those tools don't run themselves, right?
[00:02:27] They're complicated.
[00:02:28] They're big.
[00:02:28] They require a certain set of skills to operate.
[00:02:32] And so the burden now is on these data consumers.
[00:02:35] Do they know everything to operate from the data up through themselves and their consumption?
[00:02:40] No, right?
[00:02:40] They have this requirement.
[00:02:42] They're business subject matter experts.
[00:02:44] But there's a lot of – that's just the tip of the iceberg.
[00:02:47] There's a lot of the iceberg that they need to exercise to be able to get value from the data that they want to exploit.
[00:02:54] So how do we solve this today?
[00:02:56] If I can move the slides.
[00:02:58] All right, so we solve this today by pairing them with our data curators and our data engineers, those talented individuals that are technically minded enough to dig into the challenges of data and the tools that we operate.
[00:03:13] But these are incredible specialists that have a lot to work on.
[00:03:17] As a matter of fact, we all know that data is the new oil.
[00:03:20] There's data of all shapes and sizes.
[00:03:22] It comes in all forms.
[00:03:23] It is constantly changing.
[00:03:24] And so there's a huge lift, a burden on your data curators, your engineers, just to understand what data is out there and what that data means and represents and how do you think about and understand the concepts representing that data?
[00:03:40] How do we take advantage of it?
[00:03:41] That's a whole pile of knowledge they need to be able to develop on their own.
[00:03:48] And then we pair that with – now we've got to also understand the tools, right?
[00:03:52] Any vendor hyperscaler you go to these days, they're going to show you this kind of NASCAR slide, which is, you know, all of these very highly specialized tools that are built to do specific things very, very well.
[00:04:05] But now a data engineer has to be able to pick which of these tools is appropriate for their task, and then they have to know enough about that tool to actually apply it to their task.
[00:04:14] So there's even more labor and knowledge required of these engineers.
[00:04:19] So this is why it costs an arm and a leg to do these kinds of data integration, data engineering challenges, right?
[00:04:24] Now, there's a bit of a knee-jerk reaction these days, right?
[00:04:27] AI, it's new, it's powerful, it's amazing.
[00:04:31] Why don't we just feed all of our data into the AI and then ask the AI questions, right?
[00:04:35] This is kind of the naive early step in your data automation journey.
[00:04:40] This is not the right way to go about production operations.
[00:04:43] This might work when you can assume the risk of hallucinations and the costs of dealing with AI.
[00:04:50] But putting AI in your production, excuse me, in your production data flows is asking for trouble.
[00:04:57] So that's not the right solution, right?
[00:04:59] We know that AI, it hallucinates, it's non-deterministic, it's expensive to pay for the compute and the tokens and the GPUs.
[00:05:07] They're non-transparent black boxes.
[00:05:09] So if you're needing to audit work and know why things were done and what the rationale is that led to some outcome, you're not getting that from these models.
[00:05:19] And worst case, they're proprietary.
[00:05:21] You're leaving your knowledge in someone else's implementation, someone else's cloud.
[00:05:26] That's not something that you get to retain.
[00:05:28] So how do we break, how do we give you some of that power of AI to automate, be a force multiplier to our data engineers, our data curators, but we do it without these risks?
[00:05:38] That's what Clearfracture has been working on for the past several years.
[00:05:42] And we've built out Belvedere, your distinguished steward of data, to be this force multiplier to your engineers.
[00:05:51] It is a, we like to think of it as, we call it an agentic data manager, but think of it as like a data ops control plane.
[00:05:58] It is an assistant between your engineers and the tools they're already using, the tools that are already specialized, that already run at scale, that are already accredited, that already work.
[00:06:12] But we've got to be able to configure them and use them faster than we are today.
[00:06:17] So Belvedere is that control plane.
[00:06:20] It's this agentic assistant that reaches into your environment on your behalf to make sense of it, to operate it, and to accelerate the work you're doing.
[00:06:30] That's the right place for AI.
[00:06:32] The work when it runs is still your deterministic code that scales well and is affordable and is accredited and is traceable, is deterministic, right?
[00:06:40] That is what you want, but you want the AI to help you configure it and do that quickly.
[00:06:46] So that's what we've built out with Belvedere.
[00:06:49] Another way of looking at that is to imagine your environment is up in the upper right-hand corner of this picture with your data sources that you're trying to exploit, your data platforms down, and that can be in many different vendors, right?
[00:06:59] We're talking here on the slide, we've got Airflow and Glue and DBT, but you've got different platforms that you're doing that work with that you already trust.
[00:07:06] And now your data curators and engineers, instead of them going into Airflow or DBT and Glue and doing the work right there,
[00:07:15] they can step back behind Belvedere.
[00:07:17] And Belvedere really has these two arms.
[00:07:19] There's a learning arm and a working arm, a design arm.
[00:07:23] And so the learning arm allows you to point Belvedere at your environment, and it reaches in the environment and studies it autonomously.
[00:07:31] And we're going to show you some of this, right?
[00:07:32] It's going to learn what data sources you have, what are the concepts in those data sources, what are the behaviors, how are those sources being used, what are the tools and algorithms that you have to take actions against those data repositories?
[00:07:47] And it's building all of this knowledge into a knowledge base, a dynamic data catalog, a representation of your environment at the systems level.
[00:07:57] Think of it as like it's doing automatic systems engineering.
[00:08:00] It's building out component block diagrams.
[00:08:03] It's building out activity flow diagrams, lineage of all of these things touching each other to understand your environment on your behalf, taking that burden off of you so that you can come to Belvedere and say, hey, do we have data about XYZ?
[00:08:16] And it'll go tell you, you don't have to know it yourself, Belvedere is going to give you that assistance.
[00:08:20] But we're not just building knowledge for the sake of building knowledge.
[00:08:23] We want to put this knowledge to work.
[00:08:25] We want to build new activities, new business processes, new pipelines to access data and produce data products that your consumers want.
[00:08:35] That's where the other arm comes in, this design arm, the workflow building arm of Belvedere.
[00:08:40] An engineer comes to Belvedere and says, hey, I need, I don't know, a list of bad guys.
[00:08:45] I need all the revenue numbers from this region.
[00:08:47] I need SAR data and LIDAR data from Northern Africa fused into a data cube so I can build a 3D scene.
[00:08:55] You ask for the data product or you express the data product as like a ODCS, open data contract, standard kind of contract.
[00:09:02] But you can describe it with narrative and then let Belvedere build it for you in the actual tools that you have as your runtime.
[00:09:10] You want it to be built on Airflow as Python?
[00:09:12] Great.
[00:09:13] You want it to be a PySpark job on Glue?
[00:09:14] No problem.
[00:09:15] Belvedere is going to build it appropriate to your runtime and give you that code.
[00:09:20] It's kind of like a coding agent in that regard.
[00:09:23] Now, there's a lot behind Belvedere, right?
[00:09:25] This is not a trivial system.
[00:09:27] This is a collection of microservices, the multiple agents in different specializations working together.
[00:09:34] But we have built it to be appropriate for the national security mission.
[00:09:37] That means it's fully containerized, can be deployed on any Kubernetes fabric, even an air-gapped fabric.
[00:09:46] That means you can deploy Belvedere with a model garden that can host its own open weights models, or if you are in an environment where you can reach out to a hosted model, you can just point the model garden to that hosted model and take advantage of it.
[00:10:01] So we give you everything you need to operate Belvedere by itself, but we also give you all the integration points to tie into your own security services, your own model hosting services, those runtimes we talked about if you want to use a runtime that's beyond what we provide out of the box.
[00:10:19] So Belvedere gives you an awful lot that we're making very easy for you to consume, but we've built that so you can deploy it anywhere you need it deployed.
[00:10:28] I'm going to go through a demo so we don't need to talk too much about the screenshots.
[00:10:31] I'm going to talk a little bit about the use cases, though.
[00:10:33] When you think about agile data integration, there's a couple places where this really matters today.
[00:10:39] This is the one that we were dealing with a lot out in Europe, right?
[00:10:42] We had a multinational coalition fight where, you know, 28-odd nations were having to work together to provide logistics support and intelligence and other stuff to the Ukrainians.
[00:10:55] That's a lot of different, I guess, nations with lots of different systems trying to come together to produce a coherent picture of what's happening on the battlefield and picture of what's happening in the logistics supply chains.
[00:11:08] And so we had to integrate lots and lots of different sources that we didn't have control of.
[00:11:14] The Polish would show up with a new drone platform with new sensors and new protocols that we would have to engineer into our command and control systems.
[00:11:23] And that was evolution was happening all the time.
[00:11:26] And, you know, you find that our engineers that were out there were spending lots of time just trying to prepare data, get it into the format that our command and control systems could ingest.
[00:11:38] That's not just unique to Europe, right?
[00:11:40] This is all over the place.
[00:11:41] We are constantly trying to evolve our sensors, evolve our command and control systems.
[00:11:45] And we don't, soldiers don't deploy with coding skills.
[00:11:48] We don't send coders to the front line.
[00:11:50] You can't do this in the field.
[00:11:52] But with Belvedere, you can.
[00:11:54] With Belvedere, it can study the protocols coming off those sensors.
[00:11:57] They can understand the APIs of the things it's trying to feed.
[00:12:00] And when an analyst, the user at the far side comes and says, hey, I need that new, that kind of data or data from that sensor in this application.
[00:12:10] Belvedere can wire up the integration code, deploy it, and it happens very rapidly.
[00:12:18] So that's a great use case, sensor fusion, as we think about that.
[00:12:22] The next one is really where you're behind the, you're kind of the back office support for an organization.
[00:12:27] And if you're doing your business analytics, you're reporting your KPI generation.
[00:12:32] Often you're building out these data lakes with a medallion architecture where you're landing data and then you're slowly transforming it into better and better artifacts that are mission or kind of question oriented, loading those up into a gold layer.
[00:12:45] And it's very important that you understand the lineage.
[00:12:47] I mean, if you're answering a question and someone's making decisions on that answer, they want to know what was the series of steps that went from the raw data to what I'm looking at now and how reliable is the raw data?
[00:12:59] How reliable is each of those steps?
[00:13:00] I need to be able to trace that.
[00:13:02] And there's always more questions.
[00:13:04] You're constantly building new artifacts and new pipelines.
[00:13:07] That's a lot of work.
[00:13:09] It's a lot of moving parts.
[00:13:10] And to keep track of it all, especially when someone builds it and then a year later they've rotated out to the next thing and now you're responsible for trying to interpret what their code did a year ago because things have changed, right?
[00:13:21] That's hard.
[00:13:22] So with Belvedere, Belvedere wires it all together for you, keeps track of why things were done the way they were done.
[00:13:28] It maintains all of that lineage explicitly.
[00:13:31] You don't have to think about doing it yourself.
[00:13:33] It's just done for you because Belvedere runs that all for you.
[00:13:37] So that's a great use case.
[00:13:38] This last one is kind of where Clearfracture came from.
[00:13:43] This is why Clearfracture initially needed Belvedere because we were doing digital footprint analysis.
[00:13:49] You know, we're trying to hoover up the internet, trying to vacuum up the internet and say, let's make sense of it and try to build dossiers on entities of interest, right?
[00:13:59] So that's a challenge.
[00:14:02] The internet is huge.
[00:14:03] We are in control of none of it, right?
[00:14:05] I mean, we don't get to tell Facebook when it can change its APIs or any other vendor that's out there, right?
[00:14:11] They're just going to do what they need to do and we're going to try to take advantage of learning from that content.
[00:14:17] So throwing people at this problem doesn't scale.
[00:14:20] You'll never have enough engineers working for weeks, like five to six weeks per data source to try to bring it into your exploitation warehouse.
[00:14:27] You don't have enough people to get through the backlog fast enough to keep pace with the change.
[00:14:34] So that's a huge reason why Clearfracture needed the automation that Belvedere can provide to keep pace with the dynamic change that you see in publicly available information.
[00:14:47] All right.
[00:14:47] With that, we're going to jump to a demo.
[00:14:50] So give me a second to switch my screens here.
[00:14:57] Now, I will apologize.
[00:15:00] My demo is actually going to be a video recording that we're going to walk through because as fast as Belvedere is, it is still an agent.
[00:15:13] So when you ask it questions, sometimes those questions, the research you do can take single-digit minutes, sometimes a little bit longer, sometimes a little shorter.
[00:15:20] But, you know, we have a short time to do a demo here, so I want to accelerate us through some of those agentic loops and waiting for the responses from the agents.
[00:15:29] So we're going to use the video.
[00:15:30] I'll walk you through it.
[00:15:30] It's a cooking show.
[00:15:32] It's going to orient you to our application and kind of show you from start to finish how I build something.
[00:15:37] And honestly, we're going to show you how to build something that would take more than a week if I was doing it manually.
[00:15:43] And we're going to do it here.
[00:15:44] I'm going to show you how it goes in just a few minutes.
[00:15:46] We're going to have the whole thing built inside Belvedere.
[00:15:48] So I'm going to start by kind of orienting you at a high level to what Belvedere is.
[00:15:51] So hopefully, interrupt me if you guys can't see the video, but I think I've successfully started it going.
[00:15:58] All right.
[00:15:59] So this is a little explainer, right?
[00:16:01] Right now, this is your environment.
[00:16:03] On the left, you have your existing sources, processors, outputs, consumers, right?
[00:16:07] That's what you're doing today.
[00:16:09] Your data engineers, your data curators, they're reaching in the tools.
[00:16:12] They're operating them themselves, right?
[00:16:14] That's great.
[00:16:15] It works.
[00:16:15] It just doesn't work as fast as we need it to.
[00:16:18] This is what I'm saying we don't want to do.
[00:16:19] A lot of our competitors are out there like, feed me your data and then ask me questions as an AI.
[00:16:23] That's great.
[00:16:24] Sometimes it works.
[00:16:25] A lot of times it doesn't.
[00:16:26] Still black box.
[00:16:27] Your data curators are wondering what's going on in there, right?
[00:16:29] So then Belvedere comes to the table.
[00:16:31] And Belvedere says, as Belvedere says, there it is.
[00:16:37] So Belvedere shows up and it's that control layer.
[00:16:40] You're still doing the things you're doing right now.
[00:16:42] You're still building those pipelines.
[00:16:43] It's just now that Belvedere is interfacing in there, helping you get those new jobs, those changed jobs done faster, right?
[00:16:51] So with that, let's dig into the application itself.
[00:16:53] If I jump back to Belvedere for a minute, the first thing we're going to do is kind of show you, and I apologize for bouncing around the video.
[00:17:02] The first thing we're going to show you is Belvedere has multiple different arms, those different functions that it's doing.
[00:17:08] The first we're going to show you is what we call deep search.
[00:17:10] Deep search is easy to understand.
[00:17:11] Everyone here has probably worked with a chat bot before.
[00:17:14] That's what we're looking at here.
[00:17:15] So we call it deep search.
[00:17:16] You want to ask any questions, it'll go research on the web or whatever sources you point us at behind the scenes.
[00:17:21] Could be your own enterprise repositories.
[00:17:23] You ask a question, it's going to say, hey, I'm going to build this plan.
[00:17:26] Here's how I'm going to go do the research.
[00:17:28] And then when the research is done, it comes back and gives you your answer.
[00:17:31] What's nice about Belvedere is it not only gives you references, though, it also says, hey, there's follow-on things.
[00:17:36] Here's questions we think you might want to ask next to continue this research, in this case, on satellite data that might be out there.
[00:17:42] We include deep search with Belvedere for two reasons.
[00:17:45] One, it allows you to research on, hey, what data do we have that I haven't yet brought into my enterprise?
[00:17:50] Maybe there's stuff out there I need to get access to, like in this case, satellite APIs that I can download from.
[00:17:55] But secondly, it's important for Belvedere to teach itself how to operate in the environment.
[00:18:00] It might need to read product docs.
[00:18:02] It might need to read your own contextual information, your policy documents, your governance materials, to teach Belvedere how to operate in that kind of, with that context.
[00:18:13] But deep search is a chatbot, right?
[00:18:15] It's one of the functions.
[00:18:16] The next thing I want to talk about is the catalog.
[00:18:18] The catalog is actually where the knowledge is stored about what actually exists in your environment.
[00:18:24] So in this case, I've gone to our catalog tab.
[00:18:26] I'm going to show you first, when you deploy Belvedere, it's not going to hack its way through your environment.
[00:18:30] You've got to point it at the things you want to access.
[00:18:34] So you go into the settings, an administrator is going to set it up, right?
[00:18:36] You're going to connect these services.
[00:18:38] And as you notice, there's a lot of different services out there.
[00:18:41] APIs, dashboards, sorry, I'm going to rewind that a little bit.
[00:18:45] There's services up there.
[00:18:46] There's APIs, dashboards, file systems, pipelines, search indices, a lot of stuff that you can go and index.
[00:18:53] We're going to drop here into databases.
[00:18:55] When you connect a new data source, there's all kinds of prebuilt connectors.
[00:18:58] You're not having to worry a lot about how do I connect to my existing systems.
[00:19:02] But once you've connected those systems and given Belvedere the appropriate credentials to access, Belvedere is then going to autonomously connect to and study the data in those sources.
[00:19:13] So in this case, in our little environment here, we've got a bunch of different databases, pipelines, containers like S3, et cetera.
[00:19:20] And Belvedere has gone through and with our learning hive, it's studying all of this information and automatically populating this catalog.
[00:19:27] So here we're going to drill into, I'm going to riff on the social media angle for a little bit.
[00:19:31] You know, here let's go find the stuff that we have related to posts.
[00:19:34] Here I have found a posts, it looks like it's in a, I'm sorry, a S3 bucket with a bunch of folders.
[00:19:41] Now, I'm going to point out what's interesting here.
[00:19:42] If you'll see on the screen, there's an object in the center middle that says number of objects.
[00:19:47] Belvedere has crawled S3 and it's found 48 files in a folder somewhere.
[00:19:52] Maybe it's a folder structure.
[00:19:54] And it said these files aren't individual items.
[00:19:57] They're a data set.
[00:19:58] They're a collection of things that we're calling posts in this case.
[00:20:02] And it's identified them as it's built this description.
[00:20:05] It has looked and said, hey, these aren't just random, we don't know what they are files.
[00:20:10] We think they're information from the Blue Sky social media platform.
[00:20:15] They're JSON API polls from Blue Sky.
[00:20:19] And it's given me the descriptions, given me their schema.
[00:20:22] It's doing a little bit of analysis of, hey, what's useful for an analyst?
[00:20:25] What's useful for a data engineer when you're dealing with this data set of 48 files?
[00:20:30] And it's even made some of that information very explicit, right?
[00:20:33] You have an explicit schema that you can work with.
[00:20:35] There's sample data.
[00:20:36] There's all kinds of other stuff, file names, et cetera, that it's already built for you.
[00:20:40] No human in the loop.
[00:20:40] Just kind of building this stuff automatically.
[00:20:42] And you can do that for not just S3 files, you can do that for code.
[00:20:46] Here we're drilling into AWS Glue pipelines, where it's saying, hey, here's Glue pipelines, jobs that are code.
[00:20:53] And I can visualize those and understand the lineage of how this code interacts and moves data between different sources inside the environment.
[00:21:01] This is the systems model that Belvedere is building for you.
[00:21:05] And that systems model is also semantic, right?
[00:21:07] So here I'm browsing now into a glossary, a semantic layer, right?
[00:21:12] It built out an OSINT glossary where I can say, hey, here's all the concepts in OSINT.
[00:21:16] And here's the social media platforms that we understand exist out there.
[00:21:20] And now I can tie these concepts like Mastodon or Blue Sky, those social media things.
[00:21:25] I can tie them to assets.
[00:21:27] So I've gone into Mastodon.
[00:21:29] Here's the assets in my data environment that are related to the Mastodon social media service.
[00:21:35] So I can use semantic organization principles to identify the data that's relevant to me.
[00:21:41] And again, whether I'm just drilling in the actual laydown of my entities, I'm drilling into the semantic organization of those concepts.
[00:21:50] I'm just searching for things in their descriptions.
[00:21:53] All of that is now available to our Belvedere users, right?
[00:21:56] That's well-organized knowledge maintained by our agents.
[00:22:00] So it's a little bit of an orientation to the knowledge base of Belvedere.
[00:22:05] Now I'm going to show you how we put it to work.
[00:22:06] For this demo, we're going to rebuild the SpaceNet 8 challenge, right?
[00:22:11] So if anyone doesn't know SpaceNet, it's an In-Q-Tel lab that kind of runs.
[00:22:15] They tried to encourage innovation in the space domain, generally by running these contests, right?
[00:22:22] And so SpaceNet 8, that particular challenge, was a flood damage detection challenge, right?
[00:22:27] They gave you a bunch of pre-event images, a bunch of post-event images from overhead EO collection.
[00:22:34] And they said, listen, build us a model that takes this before and this after and first finds the roads, finds the buildings, finds the infrastructure, and then determines what of that infrastructure was damaged by this flood that was before and after, right?
[00:22:51] So that's what you're looking at here.
[00:22:52] This is the GitHub repo where you can go and look at the data yourself.
[00:22:56] You can see there was a challenge.
[00:22:57] We didn't participate.
[00:22:58] This is a challenge that was run before.
[00:22:59] But there's a whole bunch of models that were trained that performed in one shape or in one form or fashion that performed well in this particular challenge.
[00:23:08] So we're going to go into Belvedere and we're going to rebuild that challenge.
[00:23:11] We're going to run a bunch of new data through those models and see what happens, right?
[00:23:14] So the first thing is I've got to connect to the models themselves.
[00:23:17] I've skipped some of the boring parts.
[00:23:19] I'm not going to show you how we connect again to SageMaker.
[00:23:21] But imagine we've taken all those models, we've deployed them into SageMaker, and now we've pointed Belvedere at SageMaker and says, go, we're going to figure out what's in SageMaker.
[00:23:28] So here we are looking at SageMaker.
[00:23:30] There's nothing in there, right?
[00:23:32] We're in the catalog.
[00:23:32] There's no children.
[00:23:33] There's no description.
[00:23:35] SageMaker is just an empty connection right now.
[00:23:37] So what do we do next?
[00:23:38] We need to take our Learning Hive agents and say, okay, agents, go learn about SageMaker, what we have in there, right?
[00:23:43] So the first thing we're going to do is what's called a discovery job.
[00:23:46] These are agents that are going to just find things that are out there and learn the explicit metadata about those things.
[00:23:55] So here I'm configuring an agent to go and access SageMaker and learn about it, right?
[00:24:01] Discover the things we've got deployed.
[00:24:03] It's running.
[00:24:04] As it runs, I'm going to make a point.
[00:24:06] Everything our agents do, they do it to keep a human in the loop.
[00:24:11] They make proposals, and then you can review and accept or deny those proposals, those changes in the environment.
[00:24:17] In this case, you can see we found a whole bunch of stuff, pipelines and models in SageMaker.
[00:24:22] We've accepted those.
[00:24:23] We're going to go and accept all those changes.
[00:24:24] We're going to apply those changes.
[00:24:26] And now those changes are getting written into our knowledge base, right?
[00:24:29] They get written into the knowledge base.
[00:24:30] And now if I go back to the catalog, then I say, show me what's in the catalog now.
[00:24:35] Let's go back to that SageMaker connection.
[00:24:39] Now suddenly I have all these children.
[00:24:40] I have inputs, outputs, tarballs, batch responses.
[00:24:43] I can dig through them.
[00:24:44] And now notice each of these, though, has no description.
[00:24:47] So the next thing I'm going to do is go build what we call an enrichment job.
[00:24:51] This is going to characterize.
[00:24:53] This is where we go from just explicit metadata to implicit metadata.
[00:24:57] We're going to build new knowledge that doesn't exist yet.
[00:25:00] So we're going to point it.
[00:25:01] This is me showing you how we can go and set up this agent.
[00:25:03] We pointed at a particular data source.
[00:25:05] In this case, I'm just showing you how we can connect to all kinds of different sources,
[00:25:09] whether they're any entity that's in the catalog, we can run an agent against to try to learn more about.
[00:25:14] So here I'm playing back through and saying, hey, let's go choose a different random source.
[00:25:19] And look, it shows up.
[00:25:21] I'm going to remove that because that's not actually the SageMaker one.
[00:25:24] And then we've got SageMaker.
[00:25:25] We're going to go configure the agent itself.
[00:25:27] The agent is, you know, got a bunch of parameterization that you can apply to it.
[00:25:31] Down below, we're going to override our old changes because, again, it's a demo.
[00:25:35] I've done this once or twice.
[00:25:36] You know, you can edit a bunch of different fields on the entity.
[00:25:40] In this case, we're going to use this prompt with a bunch of inserted data to kind of build out semantic descriptions about the entities.
[00:25:49] And now it's running, right?
[00:25:50] So the agent is now behind the scenes.
[00:25:51] It's running.
[00:25:52] It's using that prompt in the LLM to reason about the sample data and the contents and maybe contextual information that our research has provided.
[00:25:59] It's given me a bunch of, again, proposals about this new descriptions.
[00:26:04] And so I'm going to go ahead and say, hey, for these four things, I can describe those data sets, those models in specific ways, right?
[00:26:11] Those pipelines.
[00:26:12] And so I can review them when I say, hey, yep, I think these are reasonable.
[00:26:16] I like what you've done.
[00:26:18] Let me accept all those changes.
[00:26:19] Once again, I can apply those, have them written back into the catalog.
[00:26:23] And now I go back to the catalog and what was before we had, which was empty information, right?
[00:26:28] We knew what was there, but we didn't have any description about it.
[00:26:30] We didn't know any of that other information.
[00:26:33] So now we're going to drill in this.
[00:26:34] Suddenly now I have descriptions.
[00:26:36] I have that semantic information has been built for me into the catalog, all driven by our agents.
[00:26:42] And if I go into one of these in detail, you can actually see that it was authored by our learning hive, right?
[00:26:47] That learning hive went, did the work, pushed it into the catalog.
[00:26:50] It's automatically building out knowledge for me.
[00:26:53] Now we don't just build knowledge for knowledge's sake.
[00:26:55] We want to operationalize it, right?
[00:26:57] We know that we're trying to build an experiment.
[00:27:00] We want to run some data through these models and see what comes out the other side.
[00:27:04] So now I'm going to build a pipeline.
[00:27:06] I come into Belvedere.
[00:27:07] I create a new pipeline, calling it my flood detection pipeline, right?
[00:27:11] I'm going to do a pipeline which processes whatever pre and post event image is.
[00:27:16] So I'm going through, I'm setting up the little description here.
[00:27:19] I say, I'm typing very slow in this case.
[00:27:22] I didn't accelerate this part of the video.
[00:27:24] We say go, right?
[00:27:25] So now I'm in my pipeline, my design agent, my design hive.
[00:27:28] And the first thing I can do is I can go to my agent.
[00:27:30] I can start asking it questions.
[00:27:31] I can say, hey, you know what?
[00:27:33] Maybe I'm so naive, I don't even know what models we have deployed inside the catalog yet.
[00:27:37] So I come and say, what models do we have that can detect flood damage?
[00:27:40] And the agent comes back and says, well, here's a couple of models.
[00:27:43] Turns out we just cataloged those a minute ago.
[00:27:45] But it's telling me and now knows.
[00:27:47] It can help me navigate my own data.
[00:27:50] I might then say, okay, I know I've got models.
[00:27:53] What data do I have that I can pass through those models to actually, you know, test them out?
[00:27:58] So I go and ask that question, right?
[00:27:59] What data sets do we have that contain the right GeoJSON polygons or imagery that I can use to run against these models, right?
[00:28:09] So then the agent comes back answering that question, right?
[00:28:12] It has searched the catalog for me.
[00:28:14] And it said, hey, I have found you've got some pre-event data, some post-event data, right?
[00:28:18] You've got stuff that is compatible with the models.
[00:28:21] You know, it tells me everything it can about it.
[00:28:24] And then ultimately I said, okay, great.
[00:28:25] I understand the models, seeing the things I think it should see in the catalog.
[00:28:30] Let's now build a pipeline.
[00:28:31] I want you to build me a pipeline that takes that pre- and post-event data imagery, passes it through the winning model of the contest, and I want to make some predictions.
[00:28:40] Let's see myself what comes out the other end.
[00:28:43] So I've tasked the agent to do that, the agent builds itself a little to-do list of things it needs to work on, right?
[00:28:49] And then it's going to work and churn through these to-dos, reasoning about the logic that needs to be applied, and it's going to build me a pipeline.
[00:28:57] So now you're going to see on the right-hand side, the canvas, which is where we visualize the activity flows of our pipelines, it starts to get built out.
[00:29:05] The agent is now making proposals about the logic of my transforms, of my pipelines, the workflow.
[00:29:12] Think of this as the flowchart of steps from start to finish that it needs to go through to do this work.
[00:29:21] The thing I'll point out here is the agent has done a lot of work for me.
[00:29:23] It's kind of reasoned about all the stuff that needs to happen, the models, the sources need to be connected to, the models that need to be connected to, where those are going to run, how they're going to run, where this code gets built out.
[00:29:34] And it's put all this into my canvas.
[00:29:36] There's a lot of detail here.
[00:29:37] Each of these nodes has a narrative description of the logic inside it.
[00:29:42] It has the data connections and sample data, so I can understand what schemas are going to be input and output to those different nodes.
[00:29:49] All of this done by the agent for me.
[00:29:52] Now, notice, though, there's no code here.
[00:29:54] This is still just narrative about what needs to happen.
[00:29:57] Still proposals, so I accept the proposals.
[00:29:59] Maybe I make some manual changes, right?
[00:30:01] I can go in here and I can change the data source.
[00:30:03] Maybe I want to say, listen, you know what?
[00:30:04] I want to add even more data.
[00:30:06] I know there's a couple of other file systems or objects in my S3 buckets that I actually want to use those as well.
[00:30:13] So maybe I drill into my training data.
[00:30:14] I'm going to go and choose a couple of, you know, outside of Louisiana.
[00:30:17] I want to choose the whole bunch, the whole set of data that I wanted to bring into this, right?
[00:30:20] Let's actually use the training data that we have, right?
[00:30:24] So you can choose and change anything that the agent has proposed, right?
[00:30:28] You're in full control.
[00:30:30] Ultimately, though, it's still just an activity flow diagram.
[00:30:33] It's abstract knowledge.
[00:30:34] Then we need to turn it into code.
[00:30:36] So we have this compile and validate section where I can actually now, with a click of a button, go from abstract logic to actual code on my target runtime.
[00:30:46] In this case, we've set the runtime as Airflow.
[00:30:48] And the first thing the agent does is it tries to validate the logic.
[00:30:51] It says, listen, you could have typed anything.
[00:30:53] You could type, hey, I'm building a transporter.
[00:30:56] The agent needs to make sure that the logic of your activity flow is everything it needs to succeed.
[00:31:01] In this case, it went and said, hey, you've got a good pipeline, but there's a couple of things you might want to improve about it, right?
[00:31:06] Edge cases you maybe didn't think about, error handling that you maybe want to think about.
[00:31:10] But then it says that, but it's a valid pipeline.
[00:31:12] So let me turn it into code, compiles it to code and deploys it to your runtime.
[00:31:17] In this case, again, Airflow.
[00:31:19] So now we're going to jump to Airflow and we're going to look at what did Belvedere send to Airflow.
[00:31:24] Again, Airflow is a different product.
[00:31:25] It could be glue.
[00:31:26] It could be Palantir.
[00:31:27] It could be anything.
[00:31:28] In this case, it's Airflow.
[00:31:30] We can go and say, okay, Airflow, show me what the compilation agent, what Belvedere's compilation agent sent your way.
[00:31:35] And here's our flood detection pipeline that we just built.
[00:31:38] And there's the code.
[00:31:39] Our agent with a push of that single button deployed code onto Airflow that is now doing the work.
[00:31:47] So now let's run that code.
[00:31:48] You can either run it in Airflow.
[00:31:49] You can go back here to Belvedere and click the button to say, trigger that pipeline, run it, right?
[00:31:53] I will say there's a bunch of other stuff.
[00:31:54] If the code didn't quite work right, you can give the compiler extra context.
[00:31:58] In this case, though, we're running in the background.
[00:32:00] And now I want to go, but you know what?
[00:32:01] This is just one model.
[00:32:03] I want to run a second model against the same data.
[00:32:05] I want to compare these different model outputs against each other.
[00:32:09] So I go back to my agent.
[00:32:10] I say, hey, go and modify this pipeline.
[00:32:13] Add another branch that tests this other model on the same data.
[00:32:17] The agent thinks for a little bit, comes back and says, okay, there you go.
[00:32:21] I've now given you a new pipeline, different branch, different model, giving me the same sort of outputs that I can use to compare against them, right?
[00:32:31] Okay, so now hopefully the running of the pipeline is done.
[00:32:35] I'm now going to show you the end results.
[00:32:37] So here we are.
[00:32:38] We're now looking at what was dropped on S3.
[00:32:40] This is the original data.
[00:32:41] We have a before, after.
[00:32:43] I'm sorry, before picture.
[00:32:44] We have the after picture, right, for one of these chips.
[00:32:47] So there's the before and after.
[00:32:49] Before, that's after.
[00:32:50] And now let's look at what the model produced.
[00:32:52] So the model, after looking at before and after, it went and produced roads.
[00:32:57] So it found and identified the roads.
[00:32:59] That's the brown lines that you can see.
[00:33:01] And then it started finding and identifying buildings.
[00:33:04] Buildings in red, those are damaged buildings from the flood.
[00:33:09] Buildings in green are the not damaged buildings, right?
[00:33:11] So now I have both the source data and now the models from that pipeline have produced my predictions, and I can start assessing those predictions for whatever they want.
[00:33:21] Okay, so with that, I'm going to stop that share.
[00:33:25] That is, that's the demo, right?
[00:33:27] So you saw right there in the few minutes of me walking through that video.
[00:33:32] Now, granted, I accelerated some of the agentic thinking steps, but you can imagine.
[00:33:36] I mean, we did this here in a 10-minute long video.
[00:33:39] I just did in a few minutes what it would have taken me quite a bit longer to do if I was trying to manually build out that exact same pipeline.
[00:33:47] That's the power of Belvedere.
[00:33:48] First, it's going to help me organize my life.
[00:33:50] It's going to help me understand what I've got in all these different sources and all these different systems.
[00:33:54] It's going to help me understand how those things are interacting.
[00:33:57] It's then going to make it really easy for me to build new things, to build new data products inside my environment.
[00:34:05] And I'm doing that all with that agentic assistance where I'm still in full control.
[00:34:09] I still can audit every decision that was made, every moment that a user went and said, okay, this is what I want to move forward with.
[00:34:16] It's all logged.
[00:34:17] That's all audited.
[00:34:18] I can actually look at the reasoning of the model tied to the decisions and changes of the humans.
[00:34:22] All of that makes it auditable.
[00:34:23] And then it produces that deterministic code that we saw that runs on my existing runtime.
[00:34:28] No rip and replace of my existing tools.
[00:34:31] If you're using Databricks, keep using Databricks.
[00:34:33] If you're using Airflow, keep using Airflow.
[00:34:35] We're just going to make it that much easier for you to make those tools do what you invested in them to do.
[00:34:42] So with that, I think that's my rapid walkthrough of what we're doing here at Clearfracture, what Belvedere is currently capable of.
[00:34:51] I'd love at that point to field any questions that you might have for me.
[00:34:55] There is one in the chat that said, why is the platform called Belvedere?
[00:35:04] Is there any background?
[00:35:06] Yeah, absolutely.
[00:35:07] So you notice we describe Belvedere as the distinguished steward of data.
[00:35:11] And if you're as old as I am and some of the other founders, you know, we remember an old sitcom from our childhoods, even before then, even of Mr. Belvedere, right?
[00:35:22] Where he was this British butler that took care of this American family that was, you know, maybe bumbling a bit through life, needed a little bit of guidance, a little assistance to make sure they're doing the things right.
[00:35:33] That's how Belvedere is helping us out, right?
[00:35:35] He is there to make sure that our data operations aren't just bumbling through life.
[00:35:42] We want them to be as well organized and as assisted as Mr. Belvedere assisted that American family.
[00:35:52] Perfect. Someone else also asked, can Belvedere connect to and consume real-time databases?
[00:35:59] Yeah, as a matter of fact, now keep in mind that what Belvedere is doing is not processing the raw data.
[00:36:06] The streams of those real-time sources aren't going into our agent and our agent is doing stuff on them all the time.
[00:36:11] What our agent does is it knows, hey, there's a real-time data stream or data source that exists and it samples that stream.
[00:36:21] To understand what kind of data is on that stream, it builds all that metadata into its knowledge base and then when you say, hey, I need you to do X, Y, Z to the data coming in on that stream, it's now going to build you the code that will operate on that stream.
[00:36:38] So the code is what handles the stream.
[00:36:40] Belvedere samples the stream to understand it, but it's building the code that's going to actually operate on the stream in production.
[00:36:48] Hopefully that makes sense.
[00:36:53] Perfect.
[00:36:54] Here's a few more in the chat.
[00:36:55] There's one that says, is Belvedere delivered as a multi-tenant SAAS, single-tenant SAAS, or can it be self-hosted slash customer managed in an air-gapped or private cloud environments?
[00:37:09] It's a great question.
[00:37:10] I will say we are still a fairly young company with Belvedere.
[00:37:14] We are not yet hosting it as a service.
[00:37:17] So SAAS is off the table.
[00:37:18] We are not yet hosting multi-tenant or single-tenant SAAS.
[00:37:23] That's not what we're doing yet.
[00:37:24] We only provide Belvedere as a delivered solution that you, the customer, deploys on your own fabric, whether that's air-gapped or not.
[00:37:33] So if you've got a cloud enclave, deploy it there.
[00:37:36] If you want to run it on-prem, deploy it there.
[00:37:38] If you want to run it on-prem on a submarine, great.
[00:37:41] All that we require is that you have a Kubernetes fabric that we can deploy Belvedere on.
[00:37:52] Okay, perfect.
[00:37:53] There's another one.
[00:37:54] It says, where is the cloud instance currently hosted and does it support deployment within AWS GovCloud?
[00:38:03] So I will say that our own development environment is in AWS, right?
[00:38:07] We spend most of our time working with Belvedere deployed on AWS.
[00:38:11] We do also have it deployed on some Spark hardware from NVIDIA.
[00:38:15] But keep in mind, Belvedere is not being hosted on your behalf.
[00:38:20] If you want to yourself deploy it on AWS, absolutely.
[00:38:24] That's what we do ourselves.
[00:38:26] We will help you if you need our assistance to get it deployed on AWS, absolutely.
[00:38:30] That is definitely one of our supported, that's where we run our own dev environment.
[00:38:35] So we know Belvedere works very well on AWS's cloud.
[00:38:42] Nice.
[00:38:42] There's another one.
[00:38:43] It says, what federal security compliance certifications does Belvedere currently hold?
[00:38:49] Yeah, that's a tough question for us, right?
[00:38:51] We are still young enough that we have yet to be accredited on any of the government networks.
[00:38:56] We are running trials right now that will, in the next six months, hopefully result in an accreditation on some top secret networks.
[00:39:04] And then, of course, the SIPR and NIPR networks for, in this case, we're working through the Army.
[00:39:09] But we are not yet an accredited solution anywhere where you can take advantage of that reciprocity.
[00:39:20] Okay, here's the next one.
[00:39:22] How does Belvedere restrict data from leaving our sovereign environment when utilizing models from Model Garden or external LLM providers?
[00:39:33] Yeah, that's responsible AI and making sure that your data remains where you want it is an important question.
[00:39:41] When you're, so with Belvedere, especially if you're using Belvedere's self-hosted models in its Model Garden, your data never leaves the enclave you put it in, right?
[00:39:52] That is, one, you control the enclave.
[00:39:54] Everything's in the enclave.
[00:39:55] So Belvedere is going to live inside that bubble that you give it.
[00:39:59] If you are pointing Belvedere, that Model Garden, to a third-party hosted model, I will tell you that there are no controls that Belvedere produces that are going to prevent that third party from, I guess, copying whatever data you send to it.
[00:40:18] That is, I guess we've made the assumption that when you deploy Belvedere, you're going to be making the right choices about the model providers that are, especially if they're hosting those models, that you trust those model providers with the data that you're sending over to them.
[00:40:34] That's not something we've built any facilitation into Belvedere to try to filter the data streams that go outside of the Belvedere umbrella.
[00:40:43] But again, we're kind of assuming that you're going to work with model providers that you trust, that are already accredited for whatever data you're going to be handling.
[00:40:51] Hopefully, that was a good answer.
[00:40:54] Perfect.
[00:40:55] The next question is, what mechanisms prevent Belvedere or underlying models from using our organizational data for model retraining or external learning?
[00:41:06] Yeah, so first I'll mention, we are not training any models with Belvedere, right?
[00:41:09] Belvedere, you deploy a foundation model and an embedding model that is pre-trained to work as the backstop to Belvedere's reasoning.
[00:41:19] But Belvedere isn't training new models unless you build a pipeline that trains new models.
[00:41:25] But again, that's on you to build that pipeline.
[00:41:27] What Belvedere does do is manage the context that gets sent to those models.
[00:41:32] So we have, if you're familiar with RAG, the retrieval augmented generation, if you're familiar with the way that that kind of system takes the data that you expose it to, vectorizes that data, and makes it available for the agent to search through it and find the relevant bits that then need to be inserted into prompts that are used to govern the reasoning that the models that receive those prompts do.
[00:41:57] That's what Belvedere is doing.
[00:41:58] That's the only thing we do, right?
[00:41:59] So we're not retraining models or fine-tuning models.
[00:42:02] But we do all that collection management, you are in full control of, right?
[00:42:06] You decide at a user level which data you want Belvedere to have access to.
[00:42:13] So if you don't want Belvedere to operate on certain data, just don't give it to Belvedere.
[00:42:17] Don't provide Belvedere that access.
[00:42:19] The power of Belvedere having these collections, though, is that you can feed it your policy and your governance materials.
[00:42:27] So if you have a security classification guide or if you have a standard operating procedures or any sorts of governance documentation, you can, as an administrator especially, give that to Belvedere and then make Belvedere use those governance materials.
[00:42:41] Even when a user who knows nothing about your policies and procedures is operating with Belvedere, they'll still be bound by that collection, even if they don't have knowledge of it, right?
[00:42:52] So hopefully that gives you a sense of how Belvedere interacts with your enterprise data.
[00:42:58] Only within the constraints that you give Belvedere when you're setting up Belvedere, you give it the credentials, you give it access to the data, that's all Belvedere is going to be able to touch.
[00:43:08] Perfect.
[00:43:10] What are some best practices within Belvedere for securing data in place and enforcing zero data retention?
[00:43:18] Wow, these are some great questions.
[00:43:19] So I'll go back to the point that Belvedere is an operations layer above your existing data tools.
[00:43:28] So Belvedere is only ever going to store metadata and maybe some sample data, right?
[00:43:34] Sometimes we're going to go and sample data and we'll store that sample data so that we can reason about it.
[00:43:39] But that is not ever restoring.
[00:43:42] It's not replicating your data.
[00:43:44] It's not moving your data by itself.
[00:43:46] You can tell Belvedere, hey, I want you to move my data from point to point B.
[00:43:49] But that's not Belvedere doing the movement.
[00:43:52] Belvedere writes the code.
[00:43:53] The code does the movement.
[00:43:55] That code lives in your existing tools.
[00:43:58] So where you really need the encryption for data at rest and the access controls for being able to access data, that's still in the same place it is today.
[00:44:07] That's in your data lakes and your databases and your data processing platforms.
[00:44:14] Belvedere isn't ever restoring that data in itself.
[00:44:18] You tell Belvedere to do things to your data in those other places and Belvedere reaches out and helps those other places make their changes.
[00:44:29] So that would go back to I was mentioning governance documents before.
[00:44:32] Some of the policies you might build into Belvedere, or if you don't build them as policies, you can ask for it in your pipelines.
[00:44:39] You just tell Belvedere, listen, when you write this pipeline, make sure you encrypt the data every time it's at rest.
[00:44:46] Make sure that you give it those instructions.
[00:44:48] And when Belvedere produces the code, it will make sure that it's abiding by those decisions.
[00:44:54] And then, of course, you can manually, if you don't trust it, you can always look at the code.
[00:44:58] You can open up in an IDE.
[00:45:00] You know, we just check the code into Git.
[00:45:02] So you just open up the IDE, point it at the repo and say, okay, I want to be that sure that it's doing it right.
[00:45:07] You can review the actual code itself.
[00:45:10] Change it right there at the code if you want it to do a different thing for you.
[00:45:13] But so hopefully that makes it clear.
[00:45:15] Belvedere is living on top of your data silos as a control layer.
[00:45:20] It is not the data layer itself.
[00:45:23] That is still what it is today.
[00:45:25] That's still your existing data layer.
[00:45:29] Perfect.
[00:45:31] We also have another question.
[00:45:33] It's, does this AI agent translate to different languages?
[00:45:37] You know, it's funny you ask because, I mean, every foundation model can do translations.
[00:45:42] So, sure, I mean, by that regard, you can ask the agent to do anything.
[00:45:46] Clear Fracture, I kind of mentioned that we came from this world of exploiting publicly available information.
[00:45:52] We actually, my team, trains translation models.
[00:45:56] We have 92 different language pairs, right?
[00:45:59] So, foreign language to English, that's a language pair.
[00:46:02] 92 different languages that we train our own translation models on.
[00:46:06] We do that intentionally because, you know, social media, short form text, there's a lot of emojis, there's a lot of idioms.
[00:46:14] Translating short form text is a little bit different than translating, you know, prose, nice narrative, newspaper articles.
[00:46:20] So, often your foundation models, your translation models built for the world of well-formed content may not perform great on social media content.
[00:46:31] So, with that, you know, Belvedere isn't built to translate the data itself, but Belvedere can plug translation models into your pipelines.
[00:46:44] Whether they're Clear Fracture's translation models or a third-party's translation model, we can plug that into your pipeline to do translations in those pipelines.
[00:46:54] If you want Belvedere to talk to you in a different language, you can absolutely say, answer this question in German.
[00:46:59] It will answer the question in German.
[00:47:00] It's not of moderate usefulness for someone who is a native German speaker.
[00:47:06] But, again, the pipelines themselves can operate on whatever language you need them to.
[00:47:14] Perfect.
[00:47:15] From a cost and ROI perspective, what have early adopters typically seen in terms of reduced engineering hours and faster time to insight?
[00:47:24] Sure.
[00:47:25] So, I'll tell you, you know, we started by using Belvedere internally, right?
[00:47:29] We have another product called Revelair that does that digital footprint analysis that's in use at Cybercom and some other places.
[00:47:36] But one of the things that, again, it took us an average of six weeks when we had a new, like, if we wanted to go after QQ or Weibo or, you know, let's pick your new social media platform that people are using, right?
[00:47:51] Every time we had that new source of data with new APIs or an API change, that would be like five, six weeks for an engineer to say, listen, oh, I got to go study that source and I've got to wire it into my, you know, the Revelair data warehouse and, you know, kind of do the right data transforms and make sure that collection's happening every, however often we need a collection to happen, right?
[00:48:11] It was a lot of work, six weeks.
[00:48:13] We actually had Belvedere re-engineer work that our human engineers had done over a series of five months, right?
[00:48:20] We took five months worth of human work and we said, Belvedere, go redo that work.
[00:48:25] And Belvedere did it all in 10 minutes.
[00:48:27] So, huge ROI internally, that's one of the reasons why we're taking Belvedere out to other people's use cases because there's not enough work for Belvedere to do on our internal use cases anymore, right?
[00:48:38] We want to give Belvedere to others so that they can accelerate their workloads as much as we have our own.
[00:48:43] That's a huge ROI from months to minutes, huge ROI.
[00:48:48] And so, we're now working with a couple of new trial customers that are finding similar results in, I kind of already mentioned that we're doing some media exploitation, right?
[00:48:59] So, we have one customer doing media exploitation, helping out investigators on certain cases, right?
[00:49:04] So, how do I, if I get a random dump of digital data, maybe that's from a cyber breach, maybe that's from a random cell phone.
[00:49:10] I've got a bunch of data.
[00:49:12] How do I make sense of it?
[00:49:13] What's in there that's relevant to my investigation?
[00:49:15] Belvedere is going to make that happen very fast.
[00:49:17] We have another customer who's looking at using us to manage contract deliverables from over 70 different contractors doing maintenance work on a certain asset.
[00:49:28] And they want that, right?
[00:49:30] So, how do I go in and take 70 different inputs and merge those into a single maintenance data store so I can build predictive models about when different parts of that asset is going to fail, right?
[00:49:40] So, that's a huge data integration challenge.
[00:49:42] So, in both cases, work that they've got teams of individuals that spend weeks on are now being able to happen much faster with Belvedere.
[00:49:51] I can't really, as those, again, we're a young company, as those trials develop into kind of finished engagements, then we'll be able to speak more about who the customers are and their specific ROI.
[00:50:09] Perfect.
[00:50:10] And how do you balance full agent autonomy with the need for human analysts and data engineers to stay accountable for final pipeline quality?
[00:50:20] Yeah, absolutely.
[00:50:21] I think, hopefully, you saw, as I was going through the demo, at every step, the agent makes proposals and the humans review the proposals and accept or reject them or modify them as appropriate.
[00:50:36] And keep in mind that it's not just those agentically derived artifacts, right?
[00:50:42] You can go behind the scenes.
[00:50:44] You can go into the code itself and make edits.
[00:50:46] You can go into the data catalog and its metadata and manually make edits.
[00:50:52] So, all of those things, whether they're inside Belvedere or outside Belvedere, the human is in full control of.
[00:50:59] The Belvedere is there to try to automate your changes and make those proposals as quickly as it can.
[00:51:06] But we still believe the human is completely responsible for the decision to accept or use any one of those proposals.
[00:51:16] And that's part of what we audit, right?
[00:51:18] I mean, every time you push that accept or reject button, you've logged in.
[00:51:22] And Belvedere knows this person who was logged in, hit the accept button on this code or this proposal or this catalog description or this pipeline activity flow node.
[00:51:35] And it knows that that was tied to this conversation history.
[00:51:39] So, we can actually look at the conversation that happened.
[00:51:41] And it's tied to this lineage inside the catalog.
[00:51:44] So, all of that is auditable inside Belvedere and it's just automatically collected.
[00:51:52] Perfect.
[00:51:53] That is all we have time for.
[00:51:55] Do you have any final closing remarks to leave with the audience?
[00:51:59] Well, I'll say that we, like I said, we're a young company.
[00:52:01] We're eager and hungry to prove the incredible value that Belvedere has brought to ourselves.
[00:52:06] And looking forward to, now that we've come out of stealth, engaging with early adopters who really want to transform and modernize their data operations.
[00:52:20] Perfect.
[00:52:21] Thank you so much.
[00:52:22] I'd like to thank our participants, as well as our speakers, for being with us today.
[00:52:26] We hope that the information you received during this webinar has been helpful.
[00:52:29] Please do not hesitate or call us and email us with any questions.