ESCROW: Learning Groups and Hierarchies from Streaming Categorical Data
Abstract
Categorical records arrive over time without supplied groups or parent relationships. A system must learn useful organization while avoiding groups that merely fit noise. ESCROW addresses this task through a data-source and industry-agnostic interface. Its central contribution is a complete conditional code that charges for groups, links, memberships, and cell owners. Each cell is encoded once, even when groups overlap. The code bounds unsupported selection under a stated no-association null; a separate result controls evidence accumulated from later blocks. Applied to COBWEB, selection raises mean membership F1 from 0.415 to 0.936 at depth one and from 0.622 to 1.000 at depth two, recovering all six reference groups and their links. It improves shallow recovery but lowers recovery at depths five through seven, and selects no groups on matched null streams. Recursive search separately achieves F1 above 0.999 in a 20-stream two-level confirmation. A semantic head shares sparse counts, and a memory head adapts trust in predictors. Matched experiments on Amazon and Wikidata compare classical, zero-shot, and few-shot LLM proposals. On Amazon, selected parent-based ESCROW saves 1.434 conditional bits per record, versus 1.032 for COBWEB and 0.705 for FISHDBC, with positive savings in all five orders. We measure structural recovery, conditional compression, and conventional prediction separately. Component studies isolate the two heads and ownership costs.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.