7 Conclusion
Most of the knowledge that makes tabular data valuable does not reside in its cells but has to be discovered and derived, imposing barriers to locating and analyzing the relevant data. In this work, we posit Open Tabular Insight Extraction (OpenTI) as the task of making this knowledge accessible to those who need it, treating the problem end-to-end, from the expressed need to presenting the extracted knowledge. We develop OpenTI as an emerging field centered on insights rather than the surface-level answers that narrower framings such as table question answering, text-to-SQL, and data analysis largely optimize for. By formulating the problem around the insight needs that users hold, the realizations that describe how analytical knowledge is extracted from tabular data, and a notion of their utility that reaches beyond literal accuracy, we establish a basis on which we conceptualize systems that address OpenTI as agents that navigate a combinatorially vast realization space under partial observability. On this basis, we recast the human side of the task under a framing of cooperative interaction and formalize the evaluation of such systems. Interleaving this conceptualization with a structured survey lets us take stock of where the field stands. We find that many of the pieces OpenTI requires already exist across communities but remain fragmented and, where present, are mostly assembled around narrow answer-centered targets. Capabilities such as data retrieval, output contextualization, cooperative interaction, and inferential and causal analysis remain thin, and evaluation methodology lags behind the systems we ultimately wish to build.
We intend the conceptualization and vocabulary developed here to give the communities working towards these ends a common ground from which to advance, making it possible to understand scattered efforts as facets of one problem and to transfer results across the conventions that have kept them apart. Realizing OpenTI will require assembling these facets into systems that operate end-to-end, yet we believe it is paramount that such systems are designed and evaluated around the insights they produce rather than accuracy alone. A result that is correct on its face can still mislead when it is interpreted outside the context that gives it meaning, as when an output accurately answers the question a user posed while missing the context for the insight they actually seek. The goal we consider most important to keep in view is therefore that OpenTI systems leave people better informed, producing results that are robust and, just as crucially, contextualized so that users arrive at factual, data-driven conclusions rather than confidently mistaken ones. Realizing this makes the cooperation between user and system a first-class concern in system design rather than a peripheral one. Pursued this way, OpenTI aims to make the knowledge held in tables as accessible as the knowledge we already retrieve from documents, democratizing access to data-driven insights in a trustworthy manner.