Two national policing datasets — 5.9 million crime records and a digital forensic evidence model — engineered as AI-Native Data Products, deployed on Teradata, and handed to four analysts that work from a map of both datasets rather than guessing their way around one. Ask in plain English; no SQL, no database account, no prior knowledge of either dataset.
Capability arrives quickly, integrated and impressive. From then on you work with your own information through their product: the model of your domain is theirs, the software is theirs, and a copy of the data sits in their store. Every layer is visible only through the layer above it.
The deeper cost is not the data but the modelling work. What your entities mean, how they relate, which rules your analysts apply, what a good answer looks like in your jurisdiction — that is institutional expertise, and in this arrangement it is expressed inside the vendor's ontology, in the vendor's platform. It accrues where it is written down.
Leaving means rebuilding knowledge you produced but do not hold.
The datasets are engineered in your warehouse to document their own meaning, following an openly published standard. The map is built from that documentation, the analysts are bound to it, and the platform running them is open source and deploys inside your estate. Method, products, map, agents, platform and model — each one open, each one yours to change.
The same modelling work lands as your objects in your database: definitions in table comments, relationships in a catalogue, rules in editable content, an ontology that exports to open standards. Your expertise is written down where you keep it — by construction, not by contract.
Leaving means keeping everything, because it was always yours.
Both routes answer questions. They differ in what you own afterwards. Two worth asking of anyone proposing to do this work for you: where does the modelling of my domain physically live, and can I read and change it without your product? The answers decide whether the expertise you invest compounds inside your organisation or outside it.
Working with national crime data normally means learning it first: which of 1,628 documented columns holds what, which codes mean which categories, which tables can honestly be joined. That work stands between a question and its answer, and it is the reason most people never ask.
This removes that step. You ask in ordinary language, and four specialist analysts — each one already versed in the data — do the rest. What makes their answers usable is not that they are fast, but that they are bounded and candid: they stay inside the data they know, they say what a figure excludes, and they decline rather than improvise when the data cannot support an answer.
Neither dataset arrived in usable shape. Both were rebuilt as AI-Native Data Products — an open design standard under which a dataset is required to carry its own documentation as data: a catalogue of every entity and column, a business glossary, a cookbook of worked queries, the design decisions taken while building it, its lineage, its quality readings, and its own verdict on whether it is fit to be trusted.
It is the reason an analyst can be dropped in front of 1,628 columns and know what they mean without a human explaining them: the dataset explains itself.
This work is already done. Both products were engineered and deployed onto a Teradata system ahead of time, and that is where they live now. You are not being asked to model anything, load anything or document anything — you are being handed the finished result. Everything from here is about reaching it.
It is also the foundation everything else on this page stands on. The analysts are not clever in spite of the data; they are useful because the products were built to explain themselves. Take that away and there is no map, no boundary and no verdict to rely on — just a model guessing at columns. The standard is the enabling work, and the rest is what it makes possible.
Six modules per product: the data itself, its meaning, its documented memory, its observability, plus search and prediction where they apply.
Tables, relationships, catalogues and access roles created as real database objects — not a description of a dataset, but the dataset itself with its documentation attached.
Every product runs its own validation and publishes the result. Both policing products currently read TRUSTED, with agent use permitted.
How many victims are female?
There are 477,466 female victims. This count includes all rows where the victim's sex (V4019) is coded as '0'. Note that this figure excludes any records where the sex was reported as missing, unknown, or not applicable.
An analyst that has to discover the data as it goes will sometimes discover the wrong thing and answer anyway. These analysts start from a map: every entity in both datasets, what each column means, how the tables join, and — visible below — the one link between the two products that was declared but never actually built.
It is not a diagram drawn for this page. It is the map itself, and you can open it in the product and watch an analyst read it.
Two policing datasets are what this example happens to be built on. The method underneath is indifferent to subject matter — the same steps apply to any source you hold, and Part 02 shows a third, entirely unrelated product running on the same deployment.
The map as it actually stands, drawn from the graph itself: two data products, four modules each, and the entities inside them. Only the Domain entities are named here — those are the ones a question about the data lands on. The full map goes three levels deeper, to every column — the 1,822 the catalogue documents, and the 2,555 actually deployed.
The FBI's national incident-based reporting for 2023 — incidents, offences, victims, offenders, arrestees and property. 5.86 million records across 1,628 documented columns.
An evidence model following the CASE and UCO standards — devices, files, accounts and the relationships between them. Synthetic data: no real case material.
How many, broken down by what, over which period — and equally “what does this field actually mean?”, answered from the datasets' own documentation.
Uderia is a platform for putting AI agents to work on enterprise data without giving up control of what they may do. Agents there are not general chatbots pointed at a database: each one is bound to a defined slice of data, checked against that boundary on every question, and answerable for what it did.
The policing analysts are packaged for it as an agent pack — a ready-made team you install in one step, complete with their knowledge of the two datasets. Installing it takes a few minutes; you do not build or configure anything.
The platform is open source, licensed under the GNU Affero General Public License v3, with a permissive commercial tier for organisations that need to keep their own modifications private. You are reading about a hosted instance because it is the quickest way to try it — but nothing here is tied to it.
Sign in at tda.uderia.com and the analysts are minutes away. Right for evaluating, for a proof of concept, and for anything not bound by data-residency rules.
The same platform, deployed inside your estate. For work under residency, classification or regulatory constraint this is not a preference — it is usually the only acceptable answer, and it is why the platform is built to be deployed rather than only subscribed to.
The on-premise case goes further than hosting location. Every layer has a local option — the language model, the embeddings, the vector and graph stores, the platform's own state. Configured that way it makes no external network calls at all: the data never leaves your network, and neither does the question asked about it. For a policing, defence or health deployment that is frequently the difference between a system that can be approved and one that cannot.
The product site — what the platform is and who it is for.
Open →The platform itself. This is where you sign in, install the analysts and ask your questions.
Open →The trust collateral — how an agent is composed, signed, bounded and judged, and how a data product earns a signature. This page lives there too.
Open →No Teradata account, no database credentials, no SQL — on the hosted instance or on your own deployment alike. The datasets are reached through a connection that carries its own restricted identity.
Any platform offering to do this work will need an answer to three questions. They are worth asking early, because the answers are difficult to change later.
All 5.9 million records remain in the warehouse. The map holds meaning — definitions, deployed names, join paths — never rows. Questions are answered by querying in place, and only the rows that answer a question come back. There is no second copy to secure, govern or explain.
The datasets' meaning exports as W3C standards — RDF, SKOS, OWL, Data Cube, PROV-O, DCAT — so the model of your domain is readable by any conformant tool. It is not a proprietary object model you can only see through one vendor's product.
Open source under AGPLv3, deployable in your own estate. The guarantees on this page are enforced by code you can read, and a turn's receipt verifies offline — with no call to the vendor who produced it.
Openness here is not a licence badge on one component. It runs the length of the stack: the method the data products were designed to, the products themselves, the model of the domain, the protocols, the platform, and the model doing the thinking. There is no layer you can only rent, and none you can only look at through the one above it.
That matters most when something needs to change — a definition that is wrong for your jurisdiction, an agent that should refuse more, a model that must run inside your own walls. The party that deploys this owns every one of those decisions.
| Layer | What it is here | What you can change |
|---|---|---|
| The method | The AI-Native Data Product standard — the foundation this rests on, published openly by Teradata under Creative Commons | Build your own products to it, or extend it — the module library and design language are public |
| The data products | Ordinary Teradata tables and views. The catalogue, glossary and cookbook are data | Everything. They are your objects in your warehouse, readable and alterable with plain SQL |
| The domain models | NIBRS from the FBI's public reporting standard; CASE and UCO, open community standards for digital evidence | Extend or replace them — neither is a vendor's private schema |
| The map | Built from the products' own catalogue, exportable as W3C RDF, SKOS, OWL, Data Cube, PROV-O and DCAT | Rebuild it from your deployment, edit it, or take it to another tool entirely |
| The analysts | Profiles, instructions, documentation and worked strategies — all editable content, not compiled behaviour | Their scope, their tone, what they refuse, what they know. Add your own |
| The protocols | MCP for tools, OpenTelemetry for traces, OAuth/OIDC/SAML for identity, Ed25519 for proof | Point them at your own tooling; the receipts verify without asking us |
| The platform | Uderia, open source under AGPLv3; the Enterprise tier adds a permissive MIT licence | Read it, patch it, run your fork on your own hardware |
| The model | Ten providers, from frontier APIs to open-weight models on your own machines | Swap it per agent; nothing here is tied to one vendor's model |
One exception, stated because you would find it anyway. The platform's own reasoning prompts — Uderia's method for planning and self-correction — are encrypted below the Enterprise tier, which lifts the licence to MIT and opens them. That is Uderia's intellectual property rather than yours, and it is the only closed layer in the stack. Everything that encodes your domain is open in every tier: the products, their catalogues and comments, the map, the analysts' instructions, their documentation and their worked strategies.
The distinction is the point. A page arguing that answers should disclose what they leave out would be a poor place to leave something out.
The alternative on offer in this market is usually the opposite arrangement: one vendor supplies the ontology, the platform and the copy of your data, and each layer is inspectable only through the layer above it. That is a faster start and a much harder exit. Nothing on this page requires you to accept that trade — and if you eventually want none of it, the products keep working, because they were always just your data, documenting itself.
Nothing in this approach is about crime data. The steps were: take a source, engineer it into a product that documents itself, build the map from the product's own catalogue, and give analysts that map and a boundary. None of those steps knows what the subject matter is.
You can see that on this very deployment. Alongside the two policing products sits a third, unrelated to this pack and to policing entirely — a commercial service-and-sales domain — built to the same standard, registered in the same catalogue, and reached through the same machinery:
| Product | Domain | Where a consumer starts |
|---|---|---|
| NIBRS_Source | National crime reporting | NIBRS_Source_Domain.Incident |
| CASE_Source | Digital forensic evidence | CASE_Source_Domain.case_object |
| A third product | Commercial service and sales | …_Domain.ServiceTicket_Current |
Three unrelated subjects, one registry, one way in. Substitute claims, telecoms, tax, clinical records, logistics or anything else you hold: the work is engineering the source into a product that explains itself — after which the analysts, the map, the boundaries and the receipts follow without being rewritten for the domain.
What you are reading, then, is an example carried far enough to be checked. The pack is real, the numbers are measured, and the pattern behind it is the part that transfers.
Wiring a capable model to a database is no longer hard, and almost every assistant platform does it. The question worth asking is what the platform does once the model starts answering — whether a wrong answer can be stopped before you read it, whether anyone can tell afterwards how good it was, and whether you could prove to a third party what happened.
Those are the capabilities Uderia is built around. Each of the following is a mechanism you can inspect, not a claim about intent.
| Capability | What is actually there |
|---|---|
| An answer can be stopped | Boundaries are enforced in the request path, not written into instructions. Every data operation is compared against the agent's declared scope, and an unsupported answer can be withheld before it reaches you rather than flagged afterwards. |
| Dishonesty is caught by shape | Six distinct failure shapes are detected deterministically and at no token cost — a claim of absence when every query failed, a zero drawn from a table never shown to hold data, a quietly coarsened answer, a total spoken from a sample, an invented identifier, a citation never actually read. |
| Every turn is graded | 56 scored dimensions, deterministic checks first and model judges only where meaning genuinely requires one — so quality is measured continuously rather than sampled by whoever happens to read the output. |
| Work can be proven, not just trusted | 11 kinds of artefact — the agent, its knowledge, its data products, its skills, its workflows — can be cryptographically signed, and a turn's chain of work exports as a receipt that verifies offline, without Uderia. |
| Nothing is locked in | 10 model providers, from frontier APIs to models running on your own hardware. Ontologies export as W3C standards (RDF, SKOS, OWL, Data Cube, PROV-O, DCAT); tools speak MCP; telemetry is OpenTelemetry-shaped; sign-in is OAuth, OIDC or SAML. |
| It can run where the data must stay | Open source under AGPLv3 and deployable inside your own estate, on-premise or in your cloud. Every layer has a local option, so an air-gapped deployment is a configuration choice rather than a different product — and you can read the code that enforces all of the above rather than take it on trust. |
A capability comparison we ran across six enterprise data and AI platforms found several of these unmatched — in particular treating an ontology as a governed, versioned product, and giving an autonomous agent an identity whose permissions are a strict subset of its owner's. The individual mechanisms above are the reason, and each can be checked rather than taken on faith.
Go to tda.uderia.com, register, and confirm the verification email. If your organisation provisions accounts centrally, ask for one instead.
Signing in lands you on the welcome screen. Press Configure Application → — the same button, renamed now that you are signed in — to enter the platform.
The analysts think with a language model. Your account already lists several — Google, OpenAI, Anthropic, OpenRouter and others — but none of them has a key yet.
Open Anatomy → LLMs and press Edit on the one you have a key for. Paste the key, press Update to save, then press Test on that card and wait for ✓ Credentials valid. One working model is enough.
Open Anatomy → MCP Servers and press + Add Server. The dialog opens on Transport Type, which defaults to SSE — switch it to HTTP first, then fill in the rest:
Open Marketplace in the sidebar. It opens on Knowledge Graphs — click across to the Agent Packs tab, where Policing Data Products is listed.
The card offers two buttons. Press Fork — it gives you your own editable copy rather than a read-only subscription — and confirm with Fork Agent Pack.
Two short questions follow, and both matter:
All four analysts, the documentation library and the knowledge graph then arrive together, already wired to both.
Open Anatomy → Agent Profiles. The list is split by the four kinds of agent — IDEATE, FOCUS, OPTIMIZE, COORDINATE — and opens on IDEATE, so the coordinator is not on screen yet. Click COORDINATE, find @POLICE, and click the star at the top-left of its card.
That one click does everything: it tests the analyst, and on success makes it your default. You will see it run the test, then the star turns gold and the card is outlined in orange. That is the confirmation.
If you also want to call the specialists directly — without going through the coordinator — switch them on while you are here. Beneath each star is a small toggle; flicking it tests that analyst and makes it selectable. @NIBRS and @CASEDF are under IDEATE, @PDOCS under FOCUS.
Click Reasoning at the top of the left sidebar to return to the chat, then New Chat to open a fresh session. Because @POLICE is now your default, you simply type — there is nothing to choose.
You get 477,466 — with the note that the figure excludes records where sex was missing, unknown or not applicable.
That single answer proves the whole chain: your model works, the connection reaches the data, the coordinator routed the question to the crime-statistics specialist, that specialist found the right column unaided, and it qualified the number instead of simply stating it. Nothing further needs configuring.
The reply says “the NIBRS analyst reports…”, and you can see that it did — but not until you open the history panel, which starts closed. Use the ‹ handle on the left edge of the conversation. Your question sits there as one conversation marked L0 and tagged @POLICE — and indented underneath it, a second conversation marked L1 and tagged @NIBRS.
That indented entry is the specialist being consulted. The coordinator opened it, put its question there, and brought the answer back. Open it and you can read that exchange in full: what was asked, what the specialist did, what it found.
Four analysts arrive with the pack. Three do the work; the fourth decides which of them should. Each is bound to something different, and that binding is what each one is worth.
Every Uderia agent is one of four classes, and the tag beside each name below says which. The class decides how an agent works, not how well — you address them all the same way.
Counts, rates, distributions and trends across 1.2 million incidents — victims, offenders, arrestees, offences, property. It writes the query for you, and it already knows what the coded values mean, so you ask in words rather than in variable names.
Reads the warehouse · holds the map · 1,508 documented columns with their coded values decoded
The evidence side: devices, files, accounts, and the links between them — exhibits, chain of custody, and the typed views the product publishes over its evidence graph. Structure you can work with before any real case material is involved.
Reads the warehouse · holds the map · CASE evidence model, synthetic by design
What something means, how it is meant to be queried, and why it is the way it is — read from the products' own glossaries, query cookbooks, design decisions, naming standards and use-case bank. It reads rather than queries, so these questions cost no database work at all and cannot be got wrong by guessing.
No connection to the data itself · 10 published documents, 104 passages
The one you talk to when you do not want to think about the other three. It routes by subject — statistics to @NIBRS, evidence to @CASEDF, meaning to @PDOCS — consults more than one when a question genuinely needs both, and keeps each contribution attributed to the analyst that produced it. Where the two products have no real join, it presents two coordinated findings and refuses to invent one.
No data connection and no map of its own — it delegates, which is exactly why it needs neither
Ask the coordinator by default. One question, and it picks. You do not need to know which analyst owns which dataset — that is the job you handed over.
Each consultation gets its own conversation. When @POLICE calls a specialist it opens a separate conversation for it and puts its question there. That is where the specialist's own work is recorded, and it stays readable afterwards.
You have to open the history to see any of this. The history panel starts closed — so by default a coordinated answer looks like one conversation with one reply, and the specialists are invisible. Open it with the ‹ handle on the left edge of the conversation. Your question then appears as an entry marked L0 with the coordinator's tag, and each specialist it consulted sits indented beneath it, marked L1 with that specialist's tag. The indent is the parent-and-child relationship, and it is the only place you can see it.
Calling one directly is opt-in. The coordinator reaches all three whatever you do. But to address one yourself — typing @NIBRS rather than letting @POLICE decide — that analyst has to be switched on, which is the optional half of step 5.
To go deeper, move into that conversation. Open the L1 entry and continue there: you are then talking to that analyst directly, with its own history, and the coordinator's thread stays clean for routing. Tagging a specialist inside the coordinator's conversation is for a one-off — it works, but every such turn is added to the coordinator's context.
The map follows the specialist. @NIBRS and @CASEDF each carry the map of both products, so it is there to open beside either of them. @PDOCS reads documents rather than data, and @POLICE holds nothing of its own — ask either for the map and there is correctly none to show.
Two exchanges, with what happened underneath each.
What does the data document about how victim sex is coded?
Answers from the datasets' own glossary, cookbook and recorded design decisions — including the noted limitation that victim demographics carry missing values, and the analysis that depends on them.
What were our quarterly sales figures?
Declines. Sales data is not part of the two policing datasets this analyst is grounded on, and it says so rather than looking elsewhere for something that resembles an answer.
Everything above is enough to use this well. What follows is for when you want to see why the answers behave the way they do — four steps, five minutes, and nothing to configure.
The map at the top of this page is not an illustration. It is the real thing, three levels shallower than what is actually in there, and you can open it.
The map belongs to the two data specialists, not to the coordinator — @POLICE delegates, so it holds no map itself.
If you have already asked something, the specialist's conversation exists — open the history panel with the ‹ handle on the left edge (it starts closed), find the L1 entry indented under your question, and carry on inside it. If you have not, start a message with @NIBRS once — that opens its conversation — and work there from then on.
The resource panel is a drawer along the top of the conversation. The small ⌄ handle centred above the messages opens it, and a row of icons appears across the width of the chat.
Click Work, open Knowledge Graph, and click the map. It opens as a panel next to the conversation and stays there while you talk — which is the point. Drag its edge to give it more room.
Put a question to the specialist while the map is on screen. It redraws to what the analyst is actually reading — here, one entity opened up with its columns fanned out around it, while the analyst worked out which of them carries the fact it was asked for.
This is the clearest view of the difference. The analyst is not searching a database hoping to recognise something useful; it is reading a map it already had, and you are watching which part it reads.
The panel shows what is in play. To walk the entire map, open Inspect from Anatomy → Knowledge Graphs: it fills the window with a sample of the whole — showing 142 nodes — and lists every kind of thing in it along the top.
Both datasets are in there together: entities, their columns, the relationships between them, and the declared-but-unusable link between the two products shown for what it is. Filter by type, search for a name, or click any node to read what the product itself says about it — the definition, the coded values, whether it holds personal data.
| Worth knowing | Why |
|---|---|
| Ask in plain language | Say “female victims”, not V4019 = '0'. The coded meanings already reach the analyst. |
| CASE data is synthetic | A 44-node evidence model for working with the structure. No real case material. |
| The two datasets do not reliably join | A link between them is declared but was never built. Treat any answer that appears to combine them with suspicion. |
| An analyst fails its test | Almost always the model: open it under Anatomy → Agent Profiles and select the model you gave a credential to in step 2. |
| A refusal is information | “Not established” and “there are none” are different claims, and these analysts keep them apart deliberately. |
Both routes below are AI-Native Data Product consumption, and both are legitimate. The standard's whole promise is that a product can be handed to an agent that has never seen it before — and the first route is that promise, exercised directly: point an agent at the product, tell it how to read the catalogue, let it work things out live. It is never stale, it needs almost no setup, and for a single product with an analyst reading every answer it is often the right choice.
The second route reads the same catalogue — just earlier, and once. The map you opened at the top is not a rival source of truth; it is the product's own entity_metadata, column_catalogue, table_relationship and access_object, projected into a form the agent holds before it plans. Nothing was modelled by hand, and nothing was invented.
That earliness is what buys the difference. An agent discovering the catalogue question by question can only be asked to stay in scope; an agent holding the map can be held to it, because the boundary is a fact the platform can check against every operation.
That approach is genuinely good. It reads the truth at the moment it asks, so it is never out of date, and it needs almost no setup. The analysts described here start from exactly the same self-describing datasets. What they add is a map built in advance — of what the data contains, what the columns mean, and how the two datasets relate — and that map is what turns instructions into limits.
Asked something the data cannot answer, does it wander off and answer anyway?
The instructions say which data to use. Nothing observes whether it complied — and a confident answer drawn from the wrong place looks exactly like a right one.
The boundary is declared by the map and compared against everything the analyst does. The answer can be withheld.
“How many victims are female?” lives in one column out of 1,628.
It has to decide to look, and choose a good search. Each attempt costs a round trip, and the choice is its own judgement.
The column, its coded meanings and its sensitivity are in hand before the analyst plans anything.
A query fails, or returns nothing, or returns only a sample. What reaches you?
A failure to retrieve and a genuine absence produce the same confident sentence.
Absence claimed without a successful query; a zero from a table never shown to hold data; a quietly coarsened answer; a total spoken from a sample.
Nothing on this page required a modelling exercise on top of the product. There is no second schema, no agent-specific ontology, no annotation pass. The map is generated — and it is generated from the modules the AI-Native Data Product standard already obliges a product to publish.
Which means the discipline pays twice. The catalogue you wrote for humans and downstream systems is, without further work, the thing that lets an agent answer safely.
| What the standard asks you to publish | What it becomes here |
|---|---|
| entity_metadata · column_catalogue | The entities and attributes on the map, with their meanings and coded values in hand before the agent plans |
| table_relationship | The join paths — including the honest record of a declared link whose target was never built |
| access_object | Which deployed object a consumer should actually query, and the boundary the platform enforces |
| Business_Glossary · Design_Decision · naming_standard | The documentation an analyst searches when the question is about meaning rather than data |
| Query_Cookbook | Proven query strategies the analysts can follow instead of deriving one from scratch |
| validation_latest · data_quality_metric · data_lineage | The product's own verdict on itself, surfaced before an agent is allowed to rely on it |
The practical consequence for a product builder: the better your product documents itself, the better these analysts are, with no work directed at them specifically. A column you describe becomes a column an agent can find. A decision you record becomes an answer it can give. A link you honestly mark as unbuilt becomes a join it refuses to fabricate.
The same holds when something is wrong. If the product publishes a mistake, that mistake is written down in one place, so it can be found and corrected there — which is precisely what a mistake living only inside a model's judgement cannot be.
An answer arrives with its working kept rather than summarised — and the last of these is the one that matters, because a query that returns nothing is otherwise dropped silently on the way to an answer.
Those deterministic checks read the record before the answer reaches you. One of them is the rule that a figure taken from a sample is not a total — the shape of mistake that is hardest to notice, because every step that produced it succeeded and the number it produces looks like an answer. By default they state what they found alongside the answer; an operator can set them to withhold it instead.
The balance sits that way round deliberately: the checks that run always are the ones that cannot have an opinion, and the model judgements are the part you switch on. Of the 56 dimensions the platform defines, 35 apply to an analyst of this kind.
That record is also how the platform improves. A weakness that has been written down can be turned into a check that holds for every agent on every turn; one that exists only as someone's memory of an odd-looking answer cannot.
A map built in advance is a copy, and a copy can fall behind what it describes. That is the trade you are making, and it is worth seeing both halves of it.
Correct something in the product — a definition, or which object a consumer should query — and a map built before that correction still carries the old answer until it is refreshed.
An agent with no map has no such weakness: it reads the truth as it asks. That is a real advantage, and it is the one you are trading away.
You give up a little freshness.
The map was derived from the product's own catalogue, so it is rebuilt from the same place. A map built here against your deployment carries its source with it and offers an Update that re-reads the product, shows you exactly what changed, and applies it.
No re-modelling, no second schema to maintain, nothing to rewrite for the domain.
One action, and the boundaries hold again.
Nearly everything on this page reduces to one question, and it is not a technical one: where does the meaning of your domain live? Not the data — the meaning. What an entity is, how it relates to another, which rules your analysts apply, what a defensible answer looks like in your jurisdiction. That is the expensive part. It is the part that is genuinely yours. And where it is written down decides everything that follows.
Three layers answer that question here. Each is open on its own terms, and each is useful alone — but what makes the combination worth having is not what each one does. It is what each one makes impossible.
The products are objects in your own warehouse, at national scale, and what an analyst may read is decided by the database account it connects as — roles and views, enforced by the system of record.
Makes impossible: reaching data the account cannot read.
An openly published standard obliges a dataset to carry its own catalogue, glossary, design decisions, lineage and verdict — as tables, in your database, readable with SQL by anyone you choose.
Makes impossible: guessing what a column means.
Open source, deployable inside your estate: it builds the map from that catalogue, holds the analysts inside it, records every turn's working and checks it before an answer is delivered.
Makes impossible: a confident answer with nothing behind it.
Take any one away and the other two lose their point. Without the warehouse, the boundary is advice. Without the standard, there is no map to be bound to — only a model reading table names. Without the platform, nobody reads the working, and a plausible wrong number ships looking exactly like every right one.
That is the differentiation, and it is worth being precise about who it is against. There are two other ways to be standing here, and both are reasonable choices that people make for good reasons.
| The closed solution e.g. Palantir |
A model wired to your database most assistant platforms |
This approach Teradata & Uderia |
|
|---|---|---|---|
| Where your domain's meaning lives | In the vendor's ontology, inside their platform | Nowhere — re-inferred from column names at every question | In your database, as tables you own |
| Who can read and change it | You can, through their product | Nobody — there is nothing written down to change | Anyone with SQL, and any tool you point at it |
| What keeps an answer in bounds | The vendor's policy engine | The prompt, and hope | A declared scope, under the database account's own limits |
| What is left after the answer | A record inside their system | A chat log | The working, checked and scored, exportable as a receipt that verifies without us |
| What leaving costs | Rebuilding knowledge you produced but do not hold | Nothing — because nothing accrued | Nothing to move. It was always yours |
The middle column is the one worth dwelling on, because it is the cheapest to start and the easiest to mistake for this. Wiring a capable model to a warehouse takes an afternoon, and it will answer you fluently on the first day. What it cannot do is tell you what it left out, refuse a question outside its remit, or show you why it believes what it said — because nothing in that arrangement ever wrote those things down.
The left column is the opposite trade: everything written down, none of it yours. The modelling work still happens — your people still do it — but it accrues inside a product you rent, and it stays there when you leave.
Capability stopped being the scarce thing.
Any competent model pointed at a database will answer you today, and answer you confidently. What is scarce is an answer you would put your name to: bounded by data you still own, produced by a system you can read, carrying its own record of how it got there — and improving, because that record is what a weakness can be turned into a check from.
Policing is the worked example. The three layers are the point, and none of them is enough alone.