SQaLe

A large realistic dataset to empower small specialised text-to-SQL models

Abstract

Frontier language models in agentic pipelines lead text-to-SQL benchmarks, at a high inference cost. Small specialised models would avoid that cost, but training them needs data that reflects the scale, semantics and structure of real databases. SQaLe is a semi-synthetic text-to-SQL dataset built on 9,259 real-world schemas from SchemaPile. Our generation pipeline extends each schema, fills it with synthetic rows, generates questions and answers them with an agent, and validates every stage by execution. The result is 1,408,056 natural-language questions paired with SQL, over schemas far larger than in any existing dataset. We train Qwen3.5-2B on SQaLe with GRPO alone, and its execution accuracy on the SQaLe test set rises from 38.7% to 66.3%. The model learns to explore the database: it finds the relevant tables in large schemas and reads the values its answer needs, which training on existing datasets elicits far less.

Why SQaLe

Small specialised models avoid the inference cost of frontier agentic pipelines. Training them needs data that looks like the databases they will be used on, and the common text-to-SQL training sets use small schemas. BIRD’s training schemas have a median of 5 tables and SynSQL’s a median of 10. A model trained on them never has to search for the tables a question needs.

SQaLe starts from 9,259 real-world schemas from SchemaPile, extends each one with tables in its own style and fills it with synthetic rows. Its schemas have a median of 113 tables and 538 columns. Questions are generated from connected groups of up to 20 tables, so a model has to find the tables a question needs before it can answer. In the record in Figure 1, the question needs 3 of the schema’s 87 tables.

The dataset

question–SQL pairs
1,408,056
real-world schemas
9,259
distinct SQL queries
176,761
median tables per schema
113
median columns per schema
538
synthetic rows
108,708,694

Each record pairs a natural-language question with a schema and a gold SQL query that has been run against the populated database. Every question exists in 8 phrasings, the original and seven rewrites in different styles, and all eight share one SQL query. The split into train and test is by schema, 95% / 5%.

The dataset is published on Hugging Face in two parts that join on schema_id. trl-lab/SQaLe-2-text-to-SQL-Queries holds the questions in all eight phrasings, the gold SQL and its result. trl-lab/SQaLe-2-text-to-SQL-Schemas holds the DDL and the generated rows of each database.

Gold SQL · the same for all 8 phrasings
Result
Schema

Figure 1. One record from the test split. The tabs switch between the eight phrasings of the question; the gold SQL and its one-row result stay the same for all of them.

SQaLe’s databases hold 108,708,694 synthetic rows. Tables hold a median of 69 rows, 95.2% of tables are populated, and 95.1% of foreign-key cells are valid after repair. Every literal in a question is grounded in the data, so a model has to read values as well as table and column names.

23.4% of SQaLe queries are nested, against 7.7% in BIRD, and 40% of the queries that join chain multiple joins, against 26% in BIRD. 3.1% of queries touch five or more tables, up to 20. No BIRD or EHRSQL query touches more than four.

How it is built

The pipeline runs in five stages, and every stage is validated by execution. Value synthesis repairs rows that break key constraints, and answering sends failed or rejected queries back to the agent.

Figure 2. The generation pipeline. Return arrows mark the repair loop in value synthesis and the two feedback loops in answering.
  1. Schema collection and extension. SchemaPile provides real-world schemas from permissively licensed GitHub repositories. A tool-using agent annotates each of the 14,597 source repositories with a domain description. An LLM then extends each schema with tables that keep its naming conventions, normalisation level and foreign-key style.
  2. Value synthesis. Tables are filled in foreign-key dependency order. For each table the LLM writes a Python function from its DDL, its original rows, the allowed foreign-key values and the domain description. Fact and junction tables get more rows and a skewed key distribution, and every table is checked for primary-key uniqueness and referential integrity.
  3. Question generation. A connected subgraph of up to 20 tables is sampled along foreign keys, together with sample rows. The LLM writes questions that need every table in it, at three difficulty levels (simple, moderate, hard), with every literal grounded in the data.
  4. Style variation. Each question is rewritten into the seven styles shown in Figure 1. All eight versions share the gold SQL.
  5. Agentic answering and judging. An agent explores the live database with tools (list_tables, describe_table, sample_rows, distinct_values, run_query) and commits with submit_sql. Execution errors go back to the agent. An LLM judge checks the result against the question and sends rejected queries back for a rewrite.1

How it compares

Figure 3 compares schema size across four datasets. The median SQaLe schema has 113 tables and 538 columns. The next largest medians are EHRSQL’s 13.5 tables and 92 columns, measured on 2 schemas. SQaLe also has the most foreign-key relations, 1,196,078 across its 9,259 schemas.

Figure 3. Median tables (open circle) and median columns (filled circle) per schema, on a log scale. The table below gives the full schema statistics; EHRSQL does not report rows per table.

Training on SQaLe

To isolate the effect of the data, we train Qwen3.5-2B with GRPO directly from the base checkpoint three times, once each on SQaLe, SynSQL-2.5M and BIRD train. We call the results MSQaLe (Qwen3.5-2B trained with GRPO on SQaLe), MSynSQL and MBIRD. Apart from the training data and the strength of a length curriculum, the base model, environment, tools, reward, steps, batch size and evaluation are identical. There are no distilled traces, no teacher and no test-time scaffolding.2

On the SQaLe test set (Figure 4), MSQaLe reaches 66.3%, which is 27.6 points above the untrained model and 9.7 points behind Qwen3.6-27B. Training on BIRD adds 15.3 points and training on SynSQL 12.0. On BIRD, MSQaLe reaches 52.3% against 54.7% for MBIRD, which was trained on BIRD itself. On EHRSQL the two are tied at 23.7%.

Execution accuracy (%) by training corpus

Figure 4. Execution accuracy of Qwen3.5-2B trained with GRPO on one corpus each, on 300 SQaLe test questions, BIRD and EHRSQL. Bold marks the best of the three trained models in each column.

Figure 5 compares cost. All models run in the same agentic harness on 100 moderate SQaLe test questions. MSQaLe reaches 52% at 224 TFLOPs per question, and Qwen3.5-27B reaches 63% at 1,710 TFLOPs. Among general-purpose models in the 9–12B range, only Qwen3.5-9B (58%) is ahead of MSQaLe.

Accuracy against model size, 100 moderate SQaLe test questions

All 25 models as a table
Figure 5. Accuracy against parameter count for every model in the same agentic harness, with dot area proportional to TFLOPs per question. Hover or focus a dot for its values; arrow keys move between dots.

How the model answers

MSQaLe never sees the schema. In each round it writes a short plan and then emits one or more tool calls. The environment runs every call, returns all the results in one message and reports how many rounds are left before the model has to answer. Figure 6 replays three of its episodes on full schemas of 106 to 239 tables, and all three end with a correct answer.3

    Reasoning, abridged

    Figure 6. Three real episodes of MSQaLe on full SQaLe schemas, with the schema withheld. Pick an episode and step through its rounds. Each round shows the model’s reasoning and every tool call it issued in that round, with the output it got back. After six rounds the environment stops the exploration and the model writes its final SQL as text. Reasoning is abridged to verbatim excerpts, identical calls within a round are shown once with a count, and long outputs are clipped.

    All three episodes start the same way. The first round lists the tables, and the next one or two rounds inspect the few tables that match the question, with four to nine distinct calls in a round. The model then drafts a query and checks its output. The first two episodes end with submit_sql. The third keeps checking until the six rounds run out, and then the environment asks for the SQL as text. Most episodes end this way: in the 108-question evaluation these episodes come from, 84 of the 108 full-schema episodes reach the round limit.

    MSQaLe also learned to batch its tool calls. The prompt asks for a single JSON object per reply, and we neither encouraged nor limited multiple calls; the environment simply runs every call it finds. During training, rounds per episode stay between 5.6 and 6.8 (Figure 7). Tool calls per episode stay near six for the first 800 steps, spike twice, and then climb steadily from step 954 to an average of 19.3 over steps 1,300 to 1,658. In the same evaluation, the untrained Qwen3.5-2B never issues more than one call per round on full schemas. MSQaLe issues several in 85% of its rounds, 5.5 calls per round of which 4.4 are distinct. The environment caps rounds, not calls, so batching lets the model inspect more of the database within the same budget. The other two GRPO runs picked up the habit less cleanly.

    Rounds and tool calls per episode during training

    Mean per reported step

    Figure 7. Rounds and tool calls per episode during GRPO training on SQaLe (run 26640986). The round limit never changes, so the growth in tool calls is growth in calls per round. The run trained for 1,800 steps; the reporting tool missed its updates after step 1,658, so the curves end there. Hover over a chart, or focus it and use the arrow keys, for the values at a step. Data: episode shape (CSV).

    What a small model learns

    The three trained models differ only in their training data, so differences in how they behave come from the data. We look at three of them.

    Schema size

    Accuracy falls for every model as the schema grows, and MSQaLe leads at every size, from 72.7% with only the gold tables to 66.3% on the full schema. With only the gold tables present, all three models open every gold table in at least 95% of episodes. On the full schema, MSQaLe opens every gold table in 89.7% of episodes, against 85.0% for MBIRD and 83.3% for MSynSQL. BIRD and SynSQL training schemas have a median of 5 and 10 tables, so models trained on them never have to search.

    Populated tables

    Before submitting, MSQaLe has seen 88% of the string literals its answer depends on. MSynSQL has seen 55%, MBIRD 40% and the untrained model 35%. SynSQL’s tables hold a median of 2 rows, and SQaLe’s hold 69.

    Schema size

    Execution accuracy (%) at three schema sizes

    String literals seen before submitting

    Share of the string literals the answer depends on that the model saw in tool output (%)

    Figure 8. Top, the three trained models at three schema sizes: the gold tables only, the gold tables plus 32 distractor tables, and the full schema. The toggle switches between accuracy and the share of episodes in which the model opened every gold table. Bottom, how many of the literals its answer needs each model has seen before it submits.

    Domains

    On SQaLe test questions whose schemas lie outside BIRD’s domains, MSQaLe leads MBIRD by 9.9 points (46.4% against 36.5%). Inside BIRD’s domains the lead is 2.8 points (39.3% against 36.5%).4

    Cite

    If you use SQaLe, please cite this paper and the earlier workshop paper.

    BibTeX

    Notes

    1. On 158 BIRD dev questions (225 judge calls), 94.5% of the queries the judge accepts are aligned with the question. The judge agrees with human labels 86.2% of the time (κ = 0.68).↩
    2. The schema is withheld. The model explores the database through tools, the five used in answering plus foreign_keys and join_path. The reward is tiered (correct result, then executes, then parses, then nothing), and partial credit never ranks a wrong query above a correct one.↩
    3. The episodes come from the full-schema condition of an earlier schema-size evaluation on 108 questions, not the 300-question evaluation behind Figures 4 and 8. Their questions use the evidence-supported phrasing, so each comes with a short evidence note.↩
    4. The domain split covers 108 schemas. That sample is too small to call the difference between the two gaps significant.↩