MortarIQStart for free

← All posts

Platforms · August 20, 2026

Is your Databricks data AI-ready? Unity Catalog can already tell you.

By the MortarIQ Founder · 6 minute read

Databricks has a distinction the other platforms in this series do not: it is usually where the AI workload itself will run. Teams pick the lakehouse because the plan is models, agents, and retrieval, not just dashboards. Which makes the readiness question unusually concrete. The pipeline that will embed your documents and the table full of unmasked emails are on the same platform, often in the same workspace, and the distance between them is one notebook.

The governance layer that can answer the question is Unity Catalog. Every catalog it governs exposes an INFORMATION_SCHEMA: tables, columns, comments, declared constraints, tags, masks, row filters, and owners, all queryable without touching a row of data. The catch, and the defining feature of Databricks estates, is the word governs. Unity Catalog only describes what lives inside it. Most estates carry a past: tables still registered in the legacy hive_metastore, which sit outside Unity Catalog’s auditing, lineage, and fine-grained access control. On Snowflake the readiness gap is usually uptake, and on BigQuery it is sprawl. On Databricks it is coverage: how much of the estate actually lives under governance at all.

Here is what an AI readiness audit reads on Databricks, the three grants it needs, and the patterns that keep showing up.

What Unity Catalog metadata answers

Everything the readiness checklist asks about a schema is sitting in catalog metadata, scored across six factors:

Contextual. Comments on tables and columns, straight from INFORMATION_SCHEMA.COLUMNS. The Databricks-specific trap is that tables here are born in code: notebooks, jobs, and pipelines create Delta tables programmatically, and programmatic creation almost never writes a COMMENT. Databricks even ships AI-generated comment suggestions to close the gap, which tells you how common the gap is. If the descriptions live in a dbt YAML file or a notebook markdown cell, the catalog is still mute, and a machine reading your catalog sees exactly what the catalog says.

Compliant.Unity Catalog’s governance surface is unusually inspectable. Tags on tables and columns say what has been classified; column masks and row filters say what is actually protected; every table records an owner. The audit measures the distance between columns that look like personal data from names and types alone and columns that actually carry a mask. With governed tags and attribute-based access control now generally available, classification can drive enforcement automatically, but a policy engine is only as good as the tagging underneath it, and tagging coverage is precisely what the metadata reveals.

Current. Last-altered timestamps for every table, read from the catalog without touching storage. Stale tables in a retrieval corpus are wrong answers delivered with confidence, and on a platform where pipelines are cheap to build and cheap to abandon, the estate accumulates tables whose jobs died quietly months ago. The catalog shows which ones, if anyone asks it.

Clean and Correlated. Databricks supports primary and foreign key declarations on Unity Catalog managed Delta tables, and does not enforce them; they are informational. They are also recent enough, reaching general availability in Databricks Runtime 15.2, that most estates were built before declaring them was possible, so almost nobody has gone back. That is measurable readiness left on the table. A declared key is the only machine-readable statement of what a row means and how tables join, and an AI consumer joining without one is guessing from column names.

Correlated, continued: lineage. This is where Databricks stands alone in this series. Unity Catalog captures lineage automatically at runtime, down to the column level, for workloads that run through it, and exposes the record in the system.access tables. That means lineage completeness is directly measurable: what fraction of your tables have any recorded lineage in the last 90 days. Snowflake exposes column-level lineage through ACCESS_HISTORY on Enterprise Edition and above, and BigQuery through its Data Lineage API once it is enabled per project; Unity Catalog captures it automatically for the workloads it governs, so on Databricks the scan reads this factor most directly. The flip side is the coverage rule again: workloads that bypass Unity Catalog, including everything still on hive_metastore, leave no trail.

Consumable. Column types and structure, from the same catalog views. A schema that leans on variant blobs where explicit types belong pushes interpretation work onto every consumer, and an AI consumer does that interpretation silently. Type discipline is visible in metadata, one column at a time.

The scan flow: a read-only connection reads catalog metadata, the six factors are scored against the selected workload profile, and a report with a fix plan comes out

See your readiness score and your biggest blocker in minutes. Read-only credentials, metadata only, starts free.

Run the free scanSee a sample report first

Three grants, no row access

A readiness audit is only worth running if the access it asks for is boring. On Databricks, a MortarIQ scan needs a workspace URL, a SQL warehouse, and a personal access token whose identity holds three grants on the catalogs in scope:

GRANT USE CATALOG ON CATALOG <catalog> TO `<token-identity>`;
GRANT USE SCHEMA  ON SCHEMA  <catalog>.<schema> TO `<token-identity>`;
GRANT BROWSE      ON CATALOG <catalog> TO `<token-identity>`;

That is catalog visibility, nothing more. The identity never receives SELECT on your tables, and the assessment never runs one against them; the queries go to INFORMATION_SCHEMA views over HTTPS via the SQL statement execution API, so a small warehouse handles the whole scan and suspends afterward. One signal is optional: the lineage check reads the system.access tables, which an account admin has to enable. If they are not enabled, that requirement reports not assessable rather than guessing, and everything else still scores. The exact SQL each connector runs is published and generated from source, so your security review is a verification, not a trust fall.

What to expect in a Databricks estate

The recurring shape is two estates wearing one name. The first is the governed lakehouse: Unity Catalog tables, owners assigned, lineage accumulating automatically, maybe tags on the obvious columns. The second is the residue of how the platform was actually adopted: hive_metastore schemas from the pre-migration era, tables materialized by notebooks nobody has opened since the person who wrote them left, and pipeline output that landed in whichever schema the job had permission to write. The governed estate would score respectably on its own. The other one is invisible to governance by construction, and it is rarely empty.

For an AI initiative the second estate is the risk that matters, because retrieval corpora and training sets get assembled from what exists, not from what is governed. The unmasked personal data is not in the tagged, owned, lineage-tracked tables; it is in the 2024 notebook output that no policy has ever applied to. A readiness scan does not fix that, but it makes the boundary between the two estates explicit, which is the difference between scoping your AI corpus and discovering it in production.

The honest boundary

A metadata scan has a ceiling on Databricks, same as anywhere: it cannot verify values. A comment can be stale, a tag can be misapplied, a fresh-looking table can be full of duplicates, and a scan that never reads rows cannot tell you otherwise; this one never reads rows by design. Value-level truth is a job for tooling that lives inside your perimeter with data access, like Databricks’ own data quality monitoring, and the two views complement rather than compete. What the readiness assessment gives you is the structural truth of the estate: what exists, what is documented, what is governed, what is fresh, what has lineage, and what your chosen workload requires that is missing. It prepares evidence. It does not certify compliance, and no tool that reads only metadata honestly can.

Gartner's 2025 AI Hype Cycle analysis (July 2025) found that 57% of organizations estimate their data is not AI-ready. On Databricks, the uncomfortable part is that the platform can already tell most of them exactly why, factor by factor, from metadata that costs nearly nothing to read.

Frequently asked questions

What access does an AI readiness scan need on Databricks?

A personal access token and a SQL warehouse, with three grants on the catalogs in scope: USE CATALOG, USE SCHEMA, and BROWSE. These expose catalog metadata and nothing else; the token identity never receives SELECT on table data. The metadata queries are lightweight, so a small SQL warehouse runs the scan and auto-suspends afterward, and every query each connector runs is published so a security team can verify the claim rather than take it on faith.

Does the scan read any rows from my Databricks tables?

No. The assessment queries Unity Catalog's INFORMATION_SCHEMA and system tables only: schema structure, comments, data types, declared constraints, last-altered timestamps, tags, column masks, row filters, table owners, and lineage records. It never runs a SELECT against your tables, and the published query list is generated from source so the boundary is verifiable.

Do I need Databricks system tables enabled to run a scan?

Only for one signal. The lineage completeness check reads system.access tables, which require an account admin to enable. If they are not enabled, that requirement reports not assessable rather than guessing, and the rest of the assessment scores normally from each catalog's INFORMATION_SCHEMA. Tables still registered in the legacy hive_metastore sit outside Unity Catalog governance, which the scan will surface as a coverage finding rather than silently skip.

Get your readiness score.

Connect read-only credentials and see your score and biggest blocker in minutes. Metadata only. Starts free.

Run the free scan

Or see a sample report on a fictional estate

This is the third platform in this series, after Snowflake and BigQuery. Want to see the output before granting anything? Read a sample readiness report built entirely from metadata. If your workspace cannot accept outside connections, the CLI runs the same metadata-only assessment from inside your network.

© MortarIQ
AboutBlogDocsFAQSecurityPrivacyTermsDPA

All product names, logos, and brands are property of their respective owners and are used for identification purposes only.