Back to main page

Methods & benchmarks

Overview of benchmarks

Overview of the analyzed benchmarks.

  • 42benchmarks
    analyzed
Setting
  • 1single table
  • bounded set
  • Nmulti-table provided
  • open corpus
Insight types
  • borderedrequires an artifact (a trained model, say), not just an insight

42 benchmarks

CorpusInputsValidation
SourceSupplement dataInsight typesFunctionOutput
AIT-QA 11610-K financial reports of major airline companiesMixedMarkup tableshierarchical column and row headers,metadata5151 Static, singleLookExact matchValue
ARCADE 106Kaggle, GitHubMixedStandalone filesnone released alongside the tables; the NL schema descriptions (column names and example cell values) used in prompts are derived from the tables themselves, and the notebook context is part of the input rather than supplementary corpus data1,0781 Static, multi-turnLookAggCharSimilarityValueTableList
Archer 68Spider (Yu et al., 2018), 20 of the 166 publicly released databasesSyntheticRelational DBsforeign-key relations and column type declarations from the database schema, plus a per-question commonsense knowledge statement for the questions that require commonsense reasoning1,042 Static, singleLookAggCFExact matchComponentValueTable
BEAVER Oracle data warehouse, MySQL databasesManualRelational DBscolumn names, column types, rows, column mapping annotation203 Static, singleLookAggExact matchStructural matchValueTable
BIRD 611Kaggle (32%), CTU Prague Relational Learning Repository (48%), open tables assembled and schema-standardized by the authors (20%)MixedRelational DBsdatabase description files (CSV) giving full schema names and value descriptions, one expert-annotated external knowledge evidence sentence per question, and the relational schemas with key relations between tables10,962 Static, singleLookAggExact matchValueTable
BIRD-INTERACT LIVESQLBENCHRelational DBshierarchical knowledge base, metadata files, database sandbox900 User simulationLookAggExact matchProgrammaticValueTable
BLADE 14Scientific publications studied in meta-analysis papers, crowd-sourced analysis studies, and statistics textbooksManualStandalone filesdataset description including details of each column, and the research question or hypothesis1881 Static, singleAssoSIExact matchStructural matchComponentValueText
BookSQL 7existing real-world accounting databases, anonymizedMixedRelational DBsforeign-key relations, primary keys100,000N Static, singleLookAggExact matchStructural matchProgram comparisonValueTableList
CausalReasoning / Benchmark 132peer-reviewed research papers, causal-inference textbooksManualStandalone filescolumn descriptions, study context1731 Static, singleAggSIIntSimilarityComponentValueText
ConDABench 1,855TidyTuesday, Kaggle notebooks, ScienceDirect open-access articlesMixedStandalone filessupporting reference code (Python) grounding each query-answer pair, and the source article each problem is derived from1,420N User simulationLookAggCharAssoPredLLM-as-judgeValueTextTableChart
CoSQL 1,020Databases inherited unchanged from Spider/SParC (originally sourced from Wikipedia, textbooks, and other web/DB sources)SyntheticRelational DBsforeign-key relations, DB schemas3,007 Static, multi-turnLookAggExact matchStructural matchSimilarityHumanTextTable
COTA KaggleMixedStandalone filescolumn meanings, value illustrations, client personas, dialog histories1,0131 User simulationLookAggCharAssoSIPredExact matchStructural matchSimilarityProgram comparisonValueTextTableListChart
CRT-QA 423WikipediaManualWeb tablestable titles1,0001 Static, singleLookAggAssoExact matchHumanText
DA-Code 1,733Kaggle, GitHub, Other Web SourcesMixedMixeddocumentation and specification files shipped alongside the data (e.g. data_standard.md, schema.yml), plus markdown, JSON and text resources; per-instance reference plans for the DA-Code-100 subset500N Static, singleLookAggCharAssoGrpSIPredExact matchSimilarityValueTextTableChart
DA-Dataset 689HangSeng Financial text-to-SQL dataset (DA-CCKS), BIRD (DA-BIRD)MixedRelational DBsdatabase schema (table names, column names, data types), natural language table and column descriptions, reference SQL queries and reference reports; relational structure inherited from the source databases735 Static, singleAggCharAssoStructural matchLLM-as-judgeText
DABstep 3AdyenSyntheticData lakecolumn descriptions, documentation files450 Static, singleLookAggCFExact matchSimilarityHumanValueTextList
DACO Spider, KaggleMixedMixedtable titles, column names, example values1,942N Static, singleLookAggCharAssoSimilarityLLM-as-judgeHumanText
DAComp - DA 418Web (base databases), enterprise SaaS schemas (analytical modeling layers via DE pipeline)SyntheticRelational DBsforeign-key relations, analytical modeling layers, data contracts100 Static, singleLookAggCharAssoLLM-as-judgeHumanTextChart
DSBench ModelOff, KaggleMixedStandalone filesimages, task descriptions, training and testing files, sample submission files540 Static, singleLookAggPredSimilarityProgrammaticLLM-as-judgeValueTable
FeTaQA 10,330Wikipedia (via ToTTo)ManualWeb tablespage title, section title, table title, highlighted table cells (denotations)10,3301 Static, singleLookAggStructural matchSimilarityHumanText
GRI-QA 204MixedMarkup tablesGRI topic, row and column indices4,089N Static, singleLookAggExact matchValueTextList
InfiAgent-DABench 67GitHubMixedStandalone files2571 Static, singleAggCharAssoSIPredExact matchValue
KaggleDBQA 17KaggleManualRelational DBsdatabase documentation, column descriptions, table names, database overview, categorical value descriptions272 Static, singleLookAggStructural matchValueTable
KramaBench 1,636scientific papers/data repositories (archaeology, astronomy, biomedical) and government/civic open-data portals (FTC, Wikipedia, NOAA, NIFC, EPA, US Census, Zillow, Kaggle) grounding published domain-expert analysesManualData lakeunstructured textual data accompanying the tabular files (scientific papers, yearly financial reports, infographics) that the tasks are grounded in; per-domain source documentation is described narratively in Appendix A rather than shipped as separate documentation files104 Static, singleLookAggAssoPredExact matchSimilarityComponentValueTextList
LongTableBench 6,800SyntheticMixedforeign-key relations5,950N Static, multi-turnLookAggSimilarityValueTextList
MMQA 3,312Spider (cross-domain relational databases spanning 138 domains), not WikipediaMixedRelational DBsforeign-key relations, primary keys, SQL query, gold answer, natural language question3,312 Static, singleLookAggExact matchStructural matchProgram comparisonLLM-as-judgeValueTextList
MT-RAIG 19,563SPIDER (relational database tables) and Open-WikiTable (Wikipedia web tables)MixedMixednone beyond the tables themselves; each table is serialized with its title and headers, while the foreign-key/topic groupings and the expanded natural-language facts were used only during benchmark construction and are not supplied to the evaluated system18,532 Static, singleLookAggCharAssoStructural matchLLM-as-judgeText
MULTIHIERTT FinTabNetMixedMarkup tablestable hierarchies, cell formats, unstructured text, reasoning programs, supporting facts10,440 Static, singleLookAggAssoExact matchSimilarityValueText
MultiTableQA 59,307Wikipedia (via HybridQA, SQA, TabFact and WikiTables)MixedWeb tablestable captions, column headers, and table metadata (contextual details, associated resources, and 3,000 example query/answer pairs injected into the metadata for demonstration)23,785 Static, singleLookAggExact matchStructural matchSimilarityProgrammaticValueList
NQ-Tables 169,898WikipediaMixedWeb tablesWikipedia page title per table, used as the table title by the retriever and reader; no other supplementary data is released11,628 Static, singleLookAggExact matchSimilarityValue
Open-WikiTable 24,680Wikipedia, via the WikiSQL and WikiTableQuestions table corporaMixedWeb tablesTable descriptions (page title, section title, caption) plus a table ID, all flattened into the retrieval passage; each question is additionally annotated with a gold SQL query alongside the textual answer.67,023 Static, singleLookAggExact matchStructural matchValue
OTT-QA 8,891WikipediaMixedWeb tablestext passages (5 million Wikipedia introduction passages), page title, page section title, section text; cell-wise hyperlinks in the training subset only4,372 Static, singleLookAggExact matchStructural matchSimilarityValue
RETQA 4,932MixedStandalone filestable captions, intent labels, slot labels20,762 Static, singleLookAggExact matchStructural matchSimilarityProgrammaticTextTable
Spider 1,056college database courses, SQL tutorial websites, online csv files, textbook examples, DatabaseAnswers data models, WikiSQL tablesSyntheticRelational DBsdatabase schemas with foreign-key relations between tables11,840 Static, singleLookAggExact matchStructural matchValueTable
Spider 2.0 418BigQuery public data, Snowflake Marketplace data and other database platforms (SQLite, DuckDB, PostgreSQL, ClickHouse); SQL queries from technical tutorials and community forums; data transformation projects from Fivetran and DBTMixedRelational DBsproject codebases and DBT/Fivetran project files, SQL dialect and external function documentation, external reference documents, database schemas with column descriptions, sampled cell values, and (for Spider 2.0) predefined answer files and data model descriptions632 Static, singleLookAggAssoExact matchProgrammaticTextTable
Tab-CQA 7,041Chinese financial reports, listed companies on Shenzhen Stock Exchange and Shanghai Stock ExchangeMixedStandalone filesconversation flow tags, answer type annotations109,0891 Static, multi-turnLookLookAggSimilarityValueText
TableBench 886Mixed8861 Static, singleLookAggAssoPredSimilarityProgrammaticLLM-as-judgeHumanTextChart
TableEval 617MixedMarkup tablestable captions, explanatory notes, adjacent text snippets2,3251 Static, multi-turnLookAggCharAssoLLM-as-judgeText
Tabular Math Word Problems (TABMWP) 37,644IXLMixedStandalone filestable titles, images, semi-structured text, gold solutions, options, units38,4311 Static, singleLookAggCharExact matchValueText
TACO 13,004Beijing smart city data service, Beijing Municipal Open Data Platform, U.S. Government's Open Data PortalMixedRelational DBsforeign-key relations, table/column descriptions, schema metadata13,000 Static, singleLookAggExact matchValueTable
TOPBench 35KaggleMixedStandalone filescolumn definitions, value distributions7791 Static, singlePredExact matchSimilarityLLM-as-judgeTextTable
WikiTableQuestions 2,108WikipediaManualWeb tables22,0331 Static, singleLookAggExact matchValueList

Reading the table

Corpus
#Tab number of tables, Source source of the tables,Curation how the tables are filled with data, Format how the tables are stored, and Supplement data what is provided alongside the tables.
Inputs
#Inst. the number of instances, Setting the data setting under which the evaluation operates. Protocol how the input is supplied to the system.
Validation
Function the validation function the benchmark uses to check the validity, and Output the output modalities it expects/supports.
Insight types
The insight types that the inputs in the benchmark ask for. The items with border indicate that the benchmark asks for an artifact (e.g., a prediction model) instead of some specific analytical knowledge.
  • Look lookup
  • Agg aggregation
  • Char characterization
  • Asso association
  • Grp grouping
  • SI statistical inference
  • Pred prediction
  • Int interventional
  • CF counterfactual