# SQaLe: a large realistic dataset to empower small specialised text-to-SQL models

Cornelius Wolff (1, 2), Daniel Gomm (1, 2), Madelon Hulsebos (2)
(1) University of Amsterdam, (2) Centrum Wiskunde & Informatica
Preprint, 2026

This is the plain-text version of the SQaLe project page at <https://trl-lab.github.io/sqale/>. The page draws its figures with JavaScript; here every figure is given as a table or a short description.

- Questions and SQL: <https://huggingface.co/datasets/trl-lab/SQaLe-2-text-to-SQL-Queries>
- Schemas and databases: <https://huggingface.co/datasets/trl-lab/SQaLe-2-text-to-SQL-Schemas>
- Model: <https://huggingface.co/trl-lab/qwen3.5-2b-grpo-sqale>
- Earlier version: *SQaLe: A Large Text-to-SQL Corpus Grounded in Real Schemas*, <https://arxiv.org/abs/2602.22223>
- Index for LLMs: <https://trl-lab.github.io/sqale/llms.txt>

## Abstract

Frontier language models in agentic pipelines lead text-to-SQL benchmarks, at a high inference cost. Small specialised models would avoid that cost, but training them needs data that reflects the scale, semantics and structure of real databases. SQaLe is a semi-synthetic text-to-SQL dataset built on 9,259 real-world schemas from SchemaPile. Our generation pipeline extends each schema, fills it with synthetic rows, generates questions and answers them with an agent, and validates every stage by execution. The result is 1,408,056 natural-language questions paired with SQL, over schemas far larger than in any existing dataset. We train Qwen3.5-2B on SQaLe with GRPO alone, and its execution accuracy on the SQaLe test set rises from 38.7% to 66.3%. The model learns to explore the database: it finds the relevant tables in large schemas and reads the values its answer needs, which training on existing datasets elicits far less.

## Why SQaLe

Small specialised models avoid the inference cost of frontier agentic pipelines. Training them needs data that looks like the databases they will be used on, and the common text-to-SQL training sets use small schemas. BIRD's training schemas have a median of 5 tables and SynSQL's a median of 10. A model trained on them never has to search for the tables a question needs.

SQaLe starts from 9,259 real-world schemas from SchemaPile, extends each one with tables in its own style and fills it with synthetic rows. Its schemas have a median of 113 tables and 538 columns. Questions are generated from connected groups of up to 20 tables, so a model has to find the tables a question needs before it can answer. In the record in Figure 1, the question needs 3 of the schema's 87 tables.

## The dataset

| | |
|---|---|
| question–SQL pairs | 1,408,056 |
| real-world schemas | 9,259 |
| distinct SQL queries | 176,761 |
| median tables per schema | 113 |
| median columns per schema | 538 |
| synthetic rows | 108,708,694 |

Each record pairs a natural-language question with a schema and a gold SQL query that has been run against the populated database. Every question exists in 8 phrasings, the original and seven rewrites in different styles, and all eight share one SQL query. The split into train and test is by schema, 95% / 5%.

The dataset is published on Hugging Face in two parts that join on `schema_id`. [trl-lab/SQaLe-2-text-to-SQL-Queries](https://huggingface.co/datasets/trl-lab/SQaLe-2-text-to-SQL-Queries) holds the questions in all eight phrasings, the gold SQL and its result. [trl-lab/SQaLe-2-text-to-SQL-Schemas](https://huggingface.co/datasets/trl-lab/SQaLe-2-text-to-SQL-Schemas) holds the DDL and the generated rows of each database.

### Figure 1. One record from the test split

`schema_013612` · 87 tables · 8,356 rows · moderate question · test split

| Phrasing | Question |
|---|---|
| Verbose (original) | Identify the names of employees who have received feedback from the employee named "Tuttle Gledhill" and determine which department they belong to, in order to assess the reach of his feedback initiatives. |
| Casual | Hey, can you pull up the names of everyone who got feedback from Tuttle Gledhill and tell me what departments they're in? Just need to see how far his feedback is spreading. |
| Structured | 1. Identify all employees who received feedback from "Tuttle Gledhill". 2. Record their full names. 3. Determine the department for each employee. 4. Use this data to evaluate the reach of his feedback initiatives. |
| Requirements list | Employees who received feedback from Tuttle Gledhill; their names; department assignments; feedback reach assessment. |
| Short and high level | List the names and departments of all employees who received feedback from Tuttle Gledhill. |
| Short and ambiguous | Track down who took feedback from Tuttle Gledhill and map out their teams. |
| Spelling and grammar mistakes | Find the names of the peeps who got feedback from Tuttle Gledhill and figure out what deparment they are in, so we can see how his feedback inititives are doing. |
| Evidence supported | Which employees received feedback from Tuttle Gledhill, and what are their departments?, evidence: Feedback records link a recipient to a sender. The sender is identified by the name 'Tuttle Gledhill'. |

Gold SQL, the same for all 8 phrasings:

```sql
SELECT DISTINCT e.first_name, e.last_name, d.d_name
FROM feedback_session fs
JOIN employee e   ON fs.feedback_to = e.id
JOIN department d ON e.department_id = d.id
WHERE fs.feedback_by = (
  SELECT id FROM employee
  WHERE first_name = 'Tuttle' AND last_name = 'Gledhill'
);
```

Result, 1 row:

| first_name | last_name | d_name |
|---|---|---|
| Kilburn | Kempster | Identity Access Management |

The schema has 87 tables. The gold SQL uses 3 of them: `feedback_session`, `employee` and `department`.

SQaLe's databases hold 108,708,694 synthetic rows. Tables hold a median of 69 rows, 95.2% of tables are populated, and 95.1% of foreign-key cells are valid after repair. Every literal in a question is grounded in the data, so a model has to read values as well as table and column names.

23.4% of SQaLe queries are nested, against 7.7% in BIRD, and 40% of the queries that join chain multiple joins, against 26% in BIRD. 3.1% of queries touch five or more tables, up to 20. No BIRD or EHRSQL query touches more than four.

## How it is built

The pipeline runs in five stages, and every stage is validated by execution. Value synthesis repairs rows that break key constraints, and answering sends failed or rejected queries back to the agent.

1. **Schema collection and extension.** SchemaPile provides real-world schemas from permissively licensed GitHub repositories. A tool-using agent annotates each of the 14,597 source repositories with a domain description. An LLM then extends each schema with tables that keep its naming conventions, normalisation level and foreign-key style.
2. **Value synthesis.** Tables are filled in foreign-key dependency order. For each table the LLM writes a Python function from its DDL, its original rows, the allowed foreign-key values and the domain description. Fact and junction tables get more rows and a skewed key distribution, and every table is checked for primary-key uniqueness and referential integrity.
3. **Question generation.** A connected subgraph of up to 20 tables is sampled along foreign keys, together with sample rows. The LLM writes questions that need every table in it, at three difficulty levels (simple, moderate, hard), with every literal grounded in the data.
4. **Style variation.** Each question is rewritten into the seven styles shown in Figure 1. All eight versions share the gold SQL.
5. **Agentic answering and judging.** An agent explores the live database with tools (`list_tables`, `describe_table`, `sample_rows`, `distinct_values`, `run_query`) and commits with `submit_sql`. Execution errors go back to the agent. An LLM judge checks the result against the question and sends rejected queries back for a rewrite.[^1]

## How it compares

Figure 3 compares schema size across four datasets. The median SQaLe schema has 113 tables and 538 columns. The next largest medians are EHRSQL's 13.5 tables and 92 columns, measured on 2 schemas. SQaLe also has the most foreign-key relations, 1,196,078 across its 9,259 schemas.

**Figure 3. Schema statistics.**

| Dataset | Schemas | Median tables | Median columns | Foreign keys | Median rows / table |
|---|--:|--:|--:|--:|--:|
| BIRD | 80 | 5 | 39 | 526 | 3,738 |
| EHRSQL | 2 | 13.5 | 92 | 34 | n/a |
| SynSQL | 16,575 | 10 | 72 | 159,547 | 2 |
| SQaLe | 9,259 | 113 | 538 | 1,196,078 | 69 |

## Training on SQaLe

To isolate the effect of the data, we train Qwen3.5-2B with GRPO directly from the base checkpoint three times, once each on SQaLe, SynSQL-2.5M and BIRD train. We call the results M_SQaLe (Qwen3.5-2B trained with GRPO on SQaLe), M_SynSQL and M_BIRD. Apart from the training data and the strength of a length curriculum, the base model, environment, tools, reward, steps, batch size and evaluation are identical. There are no distilled traces, no teacher and no test-time scaffolding.[^2]

On the SQaLe test set (Figure 4), M_SQaLe reaches 66.3%, which is 27.6 points above the untrained model and 9.7 points behind Qwen3.6-27B. Training on BIRD adds 15.3 points and training on SynSQL 12.0. On BIRD, M_SQaLe reaches 52.3% against 54.7% for M_BIRD, which was trained on BIRD itself. On EHRSQL the two are tied at 23.7%.

**Figure 4. Execution accuracy (%) by training corpus**, on 300 SQaLe test questions, BIRD and EHRSQL.

| Model | SQaLe test | BIRD | EHRSQL |
|---|--:|--:|--:|
| M_SQaLe (Qwen3.5-2B trained with GRPO on SQaLe) | **66.3** (+27.6) | 52.3 | **23.7** |
| M_BIRD (Qwen3.5-2B trained with GRPO on BIRD train) | 54.0 (+15.3) | **54.7** | **23.7** |
| M_SynSQL (Qwen3.5-2B trained with GRPO on SynSQL-2.5M) | 50.7 (+12.0) | 44.3 | 13.3 |
| Qwen3.5-2B (untrained base model) | 38.7 | 19.3 | 8.2 |
| Qwen3.6-27B (untrained, 27B parameters) | 76.0 | 69.3 | 55.0 |

Figure 5 compares cost. All models run in the same agentic harness on 100 moderate SQaLe test questions. M_SQaLe reaches 52% at 224 TFLOPs per question, and Qwen3.5-27B reaches 63% at 1,710 TFLOPs. Among general-purpose models in the 9–12B range, only Qwen3.5-9B (58%) is ahead of M_SQaLe.

**Figure 5. Accuracy against model size**, 100 moderate SQaLe test questions, all models in the same agentic harness.

| Model | Parameters (B) | Accuracy (%) | TFLOPs per question |
|---|--:|--:|--:|
| M_SQaLe | 2 | 52 | 224 |
| M_BIRD | 2 | 40 | 112 |
| M_SynSQL | 2 | 33 | 155 |
| Qwen3.5-2B | 2 | 24 | 168 |
| Granite-4.1-3B | 3 | 21 | 414 |
| Llama-3.2-3B | 3 | 10 | 535 |
| Hunyuan-4B | 4 | 36 | 448 |
| Qwen2.5-7B | 7 | 29 | 219 |
| Olmo-3-7B-Instruct | 7 | 12 | 1,001 |
| Olmo-3-7B-Think | 7 | 26 | 246 |
| Llama-3.1-8B | 8 | 30 | 693 |
| Qwen3-8B | 8 | 54 | 225 |
| GLM-4-9B | 9 | 27 | 823 |
| Qwen3.5-9B | 9 | 58 | 449 |
| Gemma-3-12B | 12 | 39 | 717 |
| Mistral-Nemo-12B | 12 | 14 | 1,239 |
| Nemotron-Nano-12B | 12 | 45 | 504 |
| Qwen3-14B | 14 | 54 | 369 |
| gpt-oss-20b | 20 | 16 | 1,052 |
| Mistral-Small-24B | 24 | 46 | 1,027 |
| Qwen3.5-27B | 27 | 63 | 1,710 |
| Qwen3-30B-A3B | 30 | 51 | 2,886 |
| Gemma-4-31B | 31 | 60 | 1,867 |
| Nemotron-Super-49B | 49 | 53 | 1,389 |
| Llama-3.3-70B | 70 | 49 | 3,154 |

## How the model answers

M_SQaLe never sees the schema. In each round it writes a short plan and then emits one or more tool calls. The environment runs every call, returns all the results in one message and reports how many rounds are left before the model has to answer. Figure 6 replays three of its episodes on full schemas of 106 to 239 tables, and all three end with a correct answer.[^3]

**Figure 6. Three real episodes of M_SQaLe on full SQaLe schemas, with the schema withheld.** Identical calls within a round are listed once.

**Shipping rates** · `schema_002386`, 239 tables · simple question · answered correctly

- Question: What are the base and final shipping rates for PRIORITY_INTL for a 48.64 unit package?
- Evidence: Base rate is the standard fee. Final rate applies weight-based adjustments. PRIORITY_INTL is an international shipping method.
- Rounds: 1 Survey (`list_tables`) · 2 Inspect (`describe_table`, `run_query`) · 3 Query (`run_query`) · 4 Submit (`submit_sql`)
- Final SQL: `SELECT sc.base_rate, sc.final_rate FROM shipping_rate_calculations sc WHERE sc.shipping_method_id = 'PRIORITY_INTL' AND sc.weight = 48.64`, which returns base_rate 23.9 and final_rate 71.51.
- What to look for: The sampled rows of `shipping_rate_calculations` show that rates are stored under method codes such as PRIORITY_INTL, while `shipping_methods` uses names and SHM- identifiers. The model filters on the code it read.

**Alert share** · `schema_018684`, 106 tables · hard question · answered correctly

- Question: What percentage of all system alert log entries were triggered by user ID 5?
- Evidence: Alert log entries are records of system notifications. User ID 5 is the specific administrator responsible for triggering the alerts.
- Rounds: 1 Survey (`list_tables`, `run_query`) · 2 Inspect (`describe_table`, `run_query`, `sample_rows`) · 3 Draft and check (`run_query`, `describe_table`) · 4 Verify (`run_query`) · 5 Submit (`submit_sql`)
- Final SQL: `SELECT (SELECT COUNT(*) FROM AlertLogs WHERE userId = 5) * 100.0 / (SELECT COUNT(*) FROM AlertLogs) AS percentage`, which returns 12.037037037037036.
- What to look for: In round 3 the model runs the percentage query together with its two parts, 26 entries for user 5 and 216 in total, and checks one against the other.

**Firmware devices** · `schema_007423`, 225 tables · moderate question · answered correctly

- Question: List device names requiring software dependencies with 'firmware' in the name, along with their total logged diagnostic request counts.
- Evidence: Device names are unique identifiers. Diagnostic requests are system logs. Firmware is a specific software component type.
- Rounds: 1 Survey (`list_tables`, `run_query`) · 2 Inspect (`describe_table`) · 3 Read values (`describe_table`, `sample_rows`, `run_query`) · 4 Count (`describe_table`, `run_query`) · 5 Draft (`describe_table`, `run_query`) · 6 Check (`run_query`) · final answer written as text
- Final SQL:

  ```sql
  SELECT d.name, (SELECT COUNT(*) FROM diagnostic_request_logs dr WHERE dr.device_id = d.uuid) as diagnostic_count
  FROM devices d
  JOIN device_software_dependencies ds ON d.uuid = ds.device_id
  WHERE ds.dependency_name LIKE '%firmware%'
  GROUP BY d.name
  ORDER BY diagnostic_count DESC
  ```

  It returns Gateway-Remote Site-742 (2), Hub-Lab-252 (1), Sensor-Remote Site-776 (0) and Node-Lab-964 (0).
- What to look for: The final query uses three of the 225 tables: it joins `devices` to `device_software_dependencies` and counts `diagnostic_request_logs` in a subquery. In round 3 a LIKE '%firmware%' query finds the dependency firmware-updater on four devices before the model writes any aggregate. Like most episodes, this one uses all six rounds and ends with the SQL written as text.

All three episodes start the same way. The first round lists the tables, and the next one or two rounds inspect the few tables that match the question, with four to nine distinct calls in a round. The model then drafts a query and checks its output. The first two episodes end with `submit_sql`. The third keeps checking until the six rounds run out, and then the environment asks for the SQL as text. Most episodes end this way: in the 108-question evaluation these episodes come from, 84 of the 108 full-schema episodes reach the round limit.

M_SQaLe also learned to batch its tool calls. The prompt asks for a single JSON object per reply, and we neither encouraged nor limited multiple calls; the environment simply runs every call it finds. During training, rounds per episode stay between 5.6 and 6.8 (Figure 7). Tool calls per episode stay near six for the first 800 steps, spike twice, and then climb steadily from step 954 to an average of 19.3 over steps 1,300 to 1,658. In the same evaluation, the untrained Qwen3.5-2B never issues more than one call per round on full schemas. M_SQaLe issues several in 85% of its rounds, 5.5 calls per round of which 4.4 are distinct. The environment caps rounds, not calls, so batching lets the model inspect more of the database within the same budget. The other two GRPO runs picked up the habit less cleanly.

**Figure 7. Rounds and tool calls per episode during GRPO training on SQaLe** (run 26640986). The round limit never changes, so the growth in tool calls is growth in calls per round. The run trained for 1,800 steps; the reporting tool missed its updates after step 1,658, so the curves end there. Data: [episode shape (CSV)](https://trl-lab.github.io/sqale/assets/training/26640986-Episode-shape.csv).

## What a small model learns

The three trained models differ only in their training data, so differences in how they behave come from the data. We look at three of them.

### Schema size

Accuracy falls for every model as the schema grows, and M_SQaLe leads at every size, from 72.7% with only the gold tables to 66.3% on the full schema. With only the gold tables present, all three models open every gold table in at least 95% of episodes. On the full schema, M_SQaLe opens every gold table in 89.7% of episodes, against 85.0% for M_BIRD and 83.3% for M_SynSQL. BIRD and SynSQL training schemas have a median of 5 and 10 tables, so models trained on them never have to search.

**Figure 8 (top). Execution accuracy (%) at three schema sizes**, with the share of episodes in which the model opened every gold table (%) in brackets.

| Tables in the database | M_SQaLe | M_BIRD | M_SynSQL |
|---|--:|--:|--:|
| Gold tables only | **72.7** (95.3) | 64.7 (98.7) | 59.3 (98.3) |
| +32 distractor tables | **68.0** (96.7) | 61.0 (90.7) | 52.3 (92.7) |
| Full schema | **66.3** (89.7) | 54.0 (85.0) | 50.7 (83.3) |

### Populated tables

Before submitting, M_SQaLe has seen 88% of the string literals its answer depends on. M_SynSQL has seen 55%, M_BIRD 40% and the untrained model 35%. SynSQL's tables hold a median of 2 rows, and SQaLe's hold 69.

**Figure 8 (bottom). String literals seen before submitting:** the share of the string literals the answer depends on that the model saw in tool output.

| Model | Literals seen (%) |
|---|--:|
| M_SQaLe | 88 |
| M_SynSQL | 55 |
| M_BIRD | 40 |
| Qwen3.5-2B | 35 |

### Domains

On SQaLe test questions whose schemas lie outside BIRD's domains, M_SQaLe leads M_BIRD by 9.9 points (46.4% against 36.5%). Inside BIRD's domains the lead is 2.8 points (39.3% against 36.5%).[^4]

## Cite

If you use SQaLe, please cite this paper and the earlier workshop paper.

```bibtex
@article{wolff2026sqale,
  title   = {SQaLe: A Large Realistic Dataset to Empower Small Specialised Text-to-SQL Models},
  author  = {Wolff, Cornelius and Gomm, Daniel and Hulsebos, Madelon},
  journal = {arXiv preprint},
  year    = {2026}
}

@article{wolff2025sqale,
  title   = {SQaLe: A Large Text-to-SQL Corpus Grounded in Real Schemas},
  author  = {Wolff, Cornelius and Gomm, Daniel and Hulsebos, Madelon},
  journal = {arXiv preprint arXiv:2602.22223},
  note    = {AI for Tabular Data workshop at EurIPS 2025},
  year    = {2025}
}
```

## Notes

[^1]: On 158 BIRD dev questions (225 judge calls), 94.5% of the queries the judge accepts are aligned with the question. The judge agrees with human labels 86.2% of the time (κ = 0.68).
[^2]: The schema is withheld. The model explores the database through tools, the five used in answering plus `foreign_keys` and `join_path`. The reward is tiered (correct result, then executes, then parses, then nothing), and partial credit never ranks a wrong query above a correct one.
[^3]: The episodes come from the full-schema condition of an earlier schema-size evaluation on 108 questions, not the 300-question evaluation behind Figures 4 and 8. Their questions use the evidence-supported phrasing, so each comes with a short evidence note.
[^4]: The domain split covers 108 schemas. That sample is too small to call the difference between the two gaps significant.

Contact: {cornelius.wolff, daniel.gomm, madelon.hulsebos}@cwi.nl. Website content MIT licensed.
