Skip to main content
Version: Next

Shared semantic metadata storage

Enablement and scope​

This host implementation supports the optional SDK metadata contract. It provides storage and cache identity. The metadata operations API adds authorized refresh, invalidation, and cache-inspection endpoints. The semantic-view editor exposes Sync metadata and Cache metadata controls, including from Explore. Both SEMANTIC_LAYERS and SEMANTIC_LAYER_METADATA_REFRESH_ENABLED remain off by default. A provider must explicitly declare support and supply the adapter and captured view token. Legacy providers retain their existing behavior.

For participating stored layers, the runtime-schema endpoint uses the bound adapter's catalog, so its choices follow refreshed metadata just like view discovery. Participation classification normalizes invalid stored provider types and configurations to the stable metadata configuration error. The runtime-schema endpoint retains its existing unknown-type response.

Before enabling, configure DISTRIBUTED_COORDINATION_CONFIG with Redis or Redis Sentinel, and SEMANTIC_LAYER_METADATA_NAMESPACE with a trusted, nonempty string or zero-argument callable returning the deployment and tenant namespace. Never derive that namespace from unvalidated request fields. The host combines it with the stored connection UUID, provider type, configuration and credentials under an HMAC using SECRET_KEY; keys do not contain raw credentials.

Enable only on a homogeneous compatible host/provider fleet. Redis rollback or restore can resurrect retired entries: change the namespace before re-enabling a recovered fleet. This cache does not provide durable cross-failover ordering.

Publication and separate invalidation​

Each successful catalog publication receives a fresh opaque token, even if the discovery JSON is unchanged. A discovery digest is not a complete upstream semantic-model revision. Hits retain the token and expiry. Failures retain the previous observation's original expiry, when it still exists. Catalog normalization preserves JSON numeric values, including decimals beyond binary floating-point precision and large or small exponents. Fresh metadata-database read failures report unavailable without driver or SQL details.

A single lease admits one writer; publication atomically compares its owner, installs the observation and releases the lease. Catalog invalidation atomically deletes both the observation and the old writer's authority. An expired or invalidated writer cannot publish over a newer observation.

Compatibility answers and query results include the token captured with their view's metadata and the view configuration. Compatibility also has an independent random generation: clearing compatibility retires all its selection variants without fetching metadata or invalidating query results. Late fills retain their old captured key. Existing query/RLS identity and selected-query force refresh remain in the query-cache path. No global key scan or upstream cache purge occurs. A SQL-backed chart's composite result key also captures participating semantic annotation sources. Async contribution tasks resolve totals using the dependent task's captured catalog: a matching entry is reused, while a different catalog requires recomputation before caching percentages. Failed totals acquisition (including a failed payload with an empty dataframe) stops the dependent query before contribution calculation or result caching. Pending participating tasks created without a serialized totals query fail closed; resubmit them after upgrading the fleet. Deploy workers before web nodes, or expect participating contribution tasks to fail until both are upgraded: an older worker cannot accept the serialized totals query sent by a newer web node.

The compatibility endpoint captures its generation before resolving the provider view. A clear during that resolution cannot relabel the endpoint's old answer with the new generation. Callers must not pre-resolve the view before this capture.

Bounds are a 30-second metadata I/O budget, a non-renewing lease capped by the owner’s remaining budget (and at most 60 seconds), a configurable catalog lifetime measured from acquisition start (300 seconds by default), and a 10-MiB serialized envelope limit. Cold readers wait and re-read within the same budget. A busy explicit refresh returns in_progress; an unknown write outcome returns indeterminate, not success or a blind retry.

Snapshot lifetime and chart-cache reuse​

An async query that outlasts its captured snapshot's lifetime may execute again when the browser reads back the completed task. The worker caches its result under the captured token, while read-back resolves the live snapshot; natural expiry rotates that token even when discovery returns identical fields. A forced-query nonce does not reuse a result under a different catalog token. This preserves the same freshness rule as an independent request: identical fields do not prove unchanged upstream definitions. Operators with long-running queries can raise SEMANTIC_LAYER_METADATA_SNAPSHOT_TTL_SECONDS to reduce this re-execution risk, trading slower metadata rediscovery for longer cache reuse.

SEMANTIC_LAYER_METADATA_SNAPSHOT_TTL_SECONDS sets the catalog lifetime and the independent compatibility-generation lifetime. It defaults to 300 seconds and accepts integer values from 1 through 2147483647; booleans, strings, zero, negative and out-of-range values fail with a configuration error. Catalog acquisition time counts against this lifetime. A discovery that consumes the entire lifetime fails with deadline and does not publish an expired snapshot. The setting does not extend the 30-second discovery budget or the writer lease. Atomic publication also caps freshness using the Redis lease's age, so transport wait cannot add time to a snapshot. This conservative anchor begins at lease installation, before acquisition. An exhausted publication fence rejects the write and preserves any previous snapshot.

Natural expiry still rotates the token, even when discovery returns identical fields: those fields need not contain the full metric definition. A longer lifetime lets chart results remain reachable longer but delays rediscovery of metadata changes; a shorter lifetime favors freshness and increases discovery and chart re-query work. This bounds reuse even when a chart has a longer cache_timeout. Explicit refresh still rotates immediately after successful publication. Hits and failures never renew snapshot expiry.

Configure the same value on all participating workers. Changes affect newly published snapshots and newly created or invalidated compatibility generations; existing entries retain their original TTL. Use the scoped invalidation controls when existing entries must expire earlier. Provider-supplied definition revisions may enable safe same-definition reuse in a future change; they are not supported by this setting.

Operation lifetime and transport​

HTTP requests establish the absolute monotonic deadline before authentication hooks. Each chart executed directly in a Celery worker gets its own 30-second metadata acquisition budget. Background dashboard export and async chart queries enter this scope before query construction; cache warm-up enters before each chart data command. Annotations, contribution totals and other nested work within that chart share its deadline and captured observations. Earlier task work or a slow preceding chart does not consume the next chart's budget. Exiting a chart restores the enclosing task state, including on failure. Inline workbook exports also give each chart its own acquisition scope; nested work shares that chart's budget. Other eager execution inside an HTTP request retains the request deadline. Celery tasks retain a fallback operation for non-chart work. Other synchronous host callers must enter metadata_operation() before access checks. A later store call never replenishes the budget; explicit worker budgets are capped at 30 seconds. The host passes that deadline explicitly to adapter.bind(store, deadline=...). The adapter passes it to store.read(fetch, deadline=...) for discovery and to store.refresh(fetch, deadline=...) for an explicit refresh. An earlier caller deadline narrows both store work and Redis transport; a deadline beyond the operation ceiling is rejected. A call never mutates the operation or another call's budget. Invalid/exhausted call deadlines fail before even cache-hit I/O. Access to an already-captured layer or view remains valid after that budget expires, so a long-running chart query does not lose its observation. Further metadata I/O still fails at the original deadline. Provider instances and views are scoped to that operation, so reusing a SQLAlchemy model in a later request cannot reuse an old provider observation. Parsed configurations are cached only within that operation and by their stored JSON text; changing the stored configuration invalidates the parsed value. Provider mutation cannot alter the cached parse. Flag-off provider construction retains its existing cache behavior.

Private Redis clients use the installed redis-py asyncio transport and one cancellation timeout per command, bounded by the operation's remaining time. This covers connection setup, Sentinel discovery and response parsing; retries are disabled. Configured CACHE_REDIS_SOCKET_TIMEOUT and CACHE_REDIS_SOCKET_CONNECT_TIMEOUT values are retained when shorter than the remaining budget, allowing Sentinel to try another node after a node timeout. For Sentinel deployments, set finite positive per-node timeouts inside the existing DISTRIBUTED_COORDINATION_CONFIG dictionary in superset_config.py:

DISTRIBUTED_COORDINATION_CONFIG.update(
CACHE_REDIS_SOCKET_TIMEOUT=1.0,
CACHE_REDIS_SOCKET_CONNECT_TIMEOUT=1.0,
)

These values are seconds; tune them for network and discovery latency. Standalone settings with those names are not read by the private metadata client. Unset or invalid values use the remaining operation budget, which can leave no time to try a second node. Each command creates a new Sentinel client and can pay the first node's timeout again. Unit tests verify timeout configuration, not live second-node failover. Cleanup supports both redis-py 5.0.0's close() and later aclose() clients, and explicitly disconnects the private Sentinel master pool. The synchronous bridge owns and closes each event loop/client, without changing shared coordinator pools. An uncancellable system DNS lookup may finish in its resolver thread after timeout. At most four system lookups can be active per process; each retains its admission slot until the actual lookup finishes, even when the command has timed out. Waiting for a slot consumes the original deadline. The cancelled command cannot connect or publish when that lookup finishes. Calling it inside an already-running asyncio loop fails explicitly; async host integrations need a synchronous worker. Provider fetches receive the same absolute deadline and must enforce it in their own transport. Monotonic values are never serialized or used to order publications.

The legacy /fetch_datasource_metadata and /datasource/get/semantic_view/<id>/ HTTP routes return the same safe discovery errors as REST boundaries: storage unavailability is HTTP 503 and deadline expiry is HTTP 504. Their existing access checks still run before view discovery.

Enablement limits​

Read authorization remains with the canonical caller policy, using its full chart, dashboard, guest-token or datasource context. Model construction must not replace those policies with a narrower datasource permission check.

The chart-context factory authorizes the full semantic context before column discovery, and query validation checks the completed context before execution. Maintenance commands require connection-management authority and revalidate the persisted principal, subject membership and stored binding before mutation. These prerequisites are implemented in the API/command layer; denied cold and warm reads are covered by zero-acquisition tests.

Production rollout still requires the default-off flag, a trusted tenant namespace, a compatible provider and fleet, bounded metadata database and Redis transports, topology/load/failover checks, and UI/live-provider acceptance. See metadata operations for the authority policy, error contract and remaining rollout gates.

Do not enable this path for deployments using semantic MCP tools or other unadapted async/CLI callers. Existing MCP handlers require a synchronous worker bridge with an explicit metadata operation before they can participate. An async caller or missing operation fails before Redis I/O. This change does not adapt those entry points.

The host's catalog publication check uses the caller's live metadata-DB connection without another pool checkout or ORM autoflush. PostgreSQL 10 or later with live READ COMMITTED isolation is required for this check. Other/unknown isolation levels, unsupported databases (including MySQL and SQLite), and pending ORM changes fail closed with unavailable (HTTP 503). This is an error response, not a silent cold read; no catalog is published and an existing unexpired snapshot is not replaced or extended.

READ COMMITTED also sees the caller's own writes. Revalidation therefore requires pg_catalog.pg_current_xact_id_if_assigned() IS NULL (PostgreSQL 13+) or pg_catalog.txid_current_if_assigned() IS NULL (10–12). These database checks accept ordinary permission exists()/count() SELECTs while detecting flushed, Core and raw-DBAPI writes, including writes rolled back to a SAVEPOINT. An assigned transaction ID is treated conservatively as uncertainty even if it came from an explicit ID request rather than a relevant write. No process-wide SQL listeners or SQL-shape classification are used.

A connection-level SAVEPOINT contains the isolation/provenance/scope reads. Statement failures roll back that nested scope without ending or flushing the caller's outer transaction. A lost connection or failed SAVEPOINT recovery can still make the connection unusable; worker cancellation remains cancellation. Distinct operator logs identify isolation, pending ORM state, unsupported provenance and SQL-read failures. The public category/status remain unavailable/503, with the neutral message “Semantic metadata is unavailable”.

The revalidation read uses the operator's existing connection/statement timeouts. It is synchronous and is not cancelled by the metadata budget; a slow metadata DB can extend elapsed request time, although publication is rejected once the budget has expired. Configure bounded DB transport/statement limits before fleet enablement. The Redis/provider deadline is not a universal HTTP or warehouse-query execution timeout.

Waiting readers poll at 50 ms and open private clients; Sentinel/TLS add connection setup work. Concurrent cold-reader load testing is an enablement gate for the intended fleet size. Live Sentinel failover is also a deployment check.

Read-only timing​

Host helpers expose CacheEntryInfo, containing entry kind/state, creation time, source observation time when known, inspection time, and finite/no-expiry/unknown expiry. Redis value and remaining TTL are read in one operation. Calculated wall-clock expiry is labeled an estimate; backend TTL remains authoritative.

Legacy query-result creation time reuses the existing UTC dttm, whose precision is one second. Missing timestamps remain unknown. Compatibility adds timing fields to the existing dictionary encoding and strips them from the public answer. Inspection does not fetch a provider catalog, initialize a generation or renew TTL.

For Redis data-cache inspection, the host caller must supply a deadline-bounded reader for the same cache server/database and resolve the key server-side. The helper retains that cache's prefix and serializer. Without such a reader it reports unsupported; it never falls back to an unbounded shared client. Other backends retain unknown expiry. A null cache reports disabled, and unavailable or undecodable entries do not become successful observations.

Catalog and compatibility maintenance/inspection require connection-management authority in the host command layer. Query-result inspection retains query/RLS access. The helpers accept internal resolved identities, not raw user cache keys. See metadata operations for the protected commands, route permissions, and API request and response shapes.

Verification​

The unit suite exercises the actual QueryContext result cache with identical discovery/name and results changing from 17 to 23, plus the existing compatibility endpoint. It also covers deadline propagation, failure categories, provider binding and non-mutating inspection.

Set SEMANTIC_METADATA_TEST_REDIS_URL to an isolated Redis endpoint to run tests/unit_tests/semantic_layers/metadata_redis_test.py. These tests use random owned prefixes, two processes, owner/follower barriers, blocked Redis I/O and concurrent entry replacement. They delete only their own keys. They do not prove a live dbt deployment, browser workflow, or Sentinel failover topology.

Chart-backed annotations​

A host chart's cache key includes the annotation source's metadata identity without constructing its provider or running discovery. It uses the view observation already captured in the operation, or peeks at the stored catalog snapshot. A current snapshot allows a warm host result to be served even if provider discovery is unavailable. On a miss, the annotation query context is prepared and authorized before the parent warehouse query, and its participating view is captured within the original discovery budget. A slow parent does not require a first annotation bind after that budget expires. Cache hits do not prepare or discover annotation providers. On a miss, the write key is recomputed after annotation acquisition so a concurrent refresh cannot store the result under a different observation. Refreshing the catalog changes the identity for subsequent operations. If the snapshot is missing, expired or unreadable, a unique key forces a miss; an unknown identity never reuses cached annotation data. If acquisition still cannot capture the keyed view, the host returns its data without persisting the unreachable result key. Flag-off and nonparticipating providers keep legacy keys.

Participating views must return a scope-qualified token issued through their operation's store. Unknown, forged or previously expired tokens from a provider's own cache are rejected as configuration errors; they cannot address derived hits. A view already captured in the operation retains its known observation after the discovery budget or shared entry expires. This does not authorize another store read or extend the discovery deadline.

Gevent request isolation​

Synchronous request greenlets run each private Redis asyncio loop in a dedicated four-thread metadata pool per native thread/hub, created lazily after first use; metadata outages do not occupy the hub pool used by its default DNS resolver. Queueing consumes the same operation deadline; expired queued work cannot begin a Redis command. Request cancellation signals the private task with a thread-safe callback, so abandoning a request does not leave an uncancelled command running. Other requests' loop state, clients, retry policy and timeouts remain independent. Ordinary native-thread callers keep the private-loop path and reject an already-running asyncio loop.

Size gevent workers for the expected concurrent metadata demand (four active commands per worker hub; excess requests queue within their deadline), and set shorter per-node socket/connect timeouts so an unhealthy Redis or Sentinel node does not consume the whole operation budget. The pool is lazy and recreated when the process or owning hub changes.