Skip to content

Iceberg lookup double-encodes user_meta, degrading result metadata to the fallback on every row #399

Description

@drernie

Observed

During live verification of the bucketless deployment (deployment PR quiltdata/deployment#2502, webhook v0.19.0 / 6be5bb6, stack tf-dev-bucketless), every package returned by the Iceberg linked-package lookup logs:

[warning] Athena metadata JSON was not an object metadata_type=str pkg_name=benchless/p2check

The lookup still works — the package matches and appears on the canvas — but the warning fires for every row on every Iceberg-path lookup, and the parsed metadata is silently discarded (see Impact).

Root cause

_build_iceberg_union_query projects the metadata column as:

json_format(CAST(m.metadata AS JSON)) AS user_meta,

m.metadata is a VARCHAR already containing a JSON document (populated from the packages-view user_meta). In Trino/Athena, CAST(varchar AS JSON) does not parse the string — it produces a JSON string value wrapping it (Trino docs: "Casting from VARCHAR to JSON does not parse the input"). json_format then serializes that, yielding a double-encoded value like "{\"entry_id\": ...}".

So in _parse_user_meta, json.loads(raw_meta) returns a Python str instead of a dict, triggering the warning and the {key: value} fallback:

def _parse_user_meta(self, raw_meta: Optional[str], *, pkg_name: str, key: str, value: str) -> Dict[str, Any]:
"""Parse metadata returned from Athena, falling back to the matched key/value."""
if not raw_meta:
return {key: value}
try:
parsed = json.loads(raw_meta)
if isinstance(parsed, dict):
return parsed
self.logger.warning(
"Athena metadata JSON was not an object",
pkg_name=pkg_name,
metadata_type=type(parsed).__name__,
)
except json.JSONDecodeError:
self.logger.warning("Failed to parse user_meta JSON", pkg_name=pkg_name)
return {key: value}

Note the query's own WHERE json_extract_scalar(m.metadata, '$.{key}') = ... works fine because json_extract_scalar does parse a VARCHAR argument — which is why filtering succeeds while the projection double-encodes.

The legacy _packages-view path is unaffected because it selects the column raw (SELECT ... user_meta, package_query.py#L322).

Impact

  • Deterministic, not traffic-dependent: every row through _find_unique_packages_in_iceberg double-encodes.
  • Canvas rendering is unaffected (it only uses bucket + pkg_name), but results.package_info[*].metadata degrades to just the matched {key: value} pair — any consumer of the full metadata gets the fallback.
  • One warning log line per row per lookup — noisy on busy stacks.

Suggested fix

Project the column raw, matching the legacy path:

m.metadata AS user_meta,

(or json_format(json_parse(m.metadata)) if normalization is desired). Optionally also make _parse_user_meta defensive: if json.loads yields a str, try decoding once more before falling back.

Repro

On any stack with QUILT_ICEBERG_DATABASE set and no package bucket: push a package with top-level user metadata {"experiment_id": "<entry display id>"}, trigger a canvas refresh, and watch the logs for the warning alongside the (successful) Found unique packages line.

🤖 Generated with Claude Code

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Fields

    No fields configured for issues without a type.

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions