Abstract
Frontier language models in agentic pipelines lead text-to-SQL benchmarks, at a high inference cost. Small specialised models would avoid that cost, but training them needs data that reflects the scale, semantics and structure of real databases. SQaLe is a semi-synthetic text-to-SQL dataset built on 9,259 real-world schemas from SchemaPile. Our generation pipeline extends each schema, fills it with synthetic rows, generates questions and answers them with an agent, and validates every stage by execution. The result is 1,408,056 natural-language questions paired with SQL, over schemas far larger than in any existing dataset. We train Qwen3.5-2B on SQaLe with GRPO alone, and its execution accuracy on the SQaLe test set rises from 38.7% to 66.3%. The model learns to explore the database: it finds the relevant tables in large schemas and reads the values its answer needs, which training on existing datasets elicits far less.
Why SQaLe
Small specialised models avoid the inference cost of frontier agentic pipelines. Training them needs data that looks like the databases they will be used on, and the common text-to-SQL training sets use small schemas. BIRD’s training schemas have a median of 5 tables and SynSQL’s a median of 10. A model trained on them never has to search for the tables a question needs.
SQaLe starts from 9,259 real-world schemas from SchemaPile, extends each one with tables in its own style and fills it with synthetic rows. Its schemas have a median of 113 tables and 538 columns. Questions are generated from connected groups of up to 20 tables, so a model has to find the tables a question needs before it can answer. In the record in Figure 1, the question needs 3 of the schema’s 87 tables.
The dataset
- question–SQL pairs
- 1,408,056
- real-world schemas
- 9,259
- distinct SQL queries
- 176,761
- median tables per schema
- 113
- median columns per schema
- 538
- synthetic rows
- 108,708,694
Each record pairs a natural-language question with a schema and a gold SQL query that has been run against the populated database. Every question exists in 8 phrasings, the original and seven rewrites in different styles, and all eight share one SQL query. The split into train and test is by schema, 95% / 5%.
The dataset is published on Hugging Face in two parts that join on schema_id. trl-lab/SQaLe-2-text-to-SQL-Queries holds the questions in all eight phrasings, the gold SQL and its result. trl-lab/SQaLe-2-text-to-SQL-Schemas holds the DDL and the generated rows of each database.
SQaLe’s databases hold 108,708,694 synthetic rows. Tables hold a median of 69 rows, 95.2% of tables are populated, and 95.1% of foreign-key cells are valid after repair. Every literal in a question is grounded in the data, so a model has to read values as well as table and column names.
23.4% of SQaLe queries are nested, against 7.7% in BIRD, and 40% of the queries that join chain multiple joins, against 26% in BIRD. 3.1% of queries touch five or more tables, up to 20. No BIRD or EHRSQL query touches more than four.
How it is built
The pipeline runs in five stages, and every stage is validated by execution. Value synthesis repairs rows that break key constraints, and answering sends failed or rejected queries back to the agent.
- Schema collection and extension. SchemaPile provides real-world schemas from permissively licensed GitHub repositories. A tool-using agent annotates each of the 14,597 source repositories with a domain description. An LLM then extends each schema with tables that keep its naming conventions, normalisation level and foreign-key style.
- Value synthesis. Tables are filled in foreign-key dependency order. For each table the LLM writes a Python function from its DDL, its original rows, the allowed foreign-key values and the domain description. Fact and junction tables get more rows and a skewed key distribution, and every table is checked for primary-key uniqueness and referential integrity.
- Question generation. A connected subgraph of up to 20 tables is sampled along foreign keys, together with sample rows. The LLM writes questions that need every table in it, at three difficulty levels (simple, moderate, hard), with every literal grounded in the data.
- Style variation. Each question is rewritten into the seven styles shown in Figure 1. All eight versions share the gold SQL.
- Agentic answering and judging. An agent explores the live database with tools (
list_tables,describe_table,sample_rows,distinct_values,run_query) and commits withsubmit_sql. Execution errors go back to the agent. An LLM judge checks the result against the question and sends rejected queries back for a rewrite.1
How it compares
Figure 3 compares schema size across four datasets. The median SQaLe schema has 113 tables and 538 columns. The next largest medians are EHRSQL’s 13.5 tables and 92 columns, measured on 2 schemas. SQaLe also has the most foreign-key relations, 1,196,078 across its 9,259 schemas.
Training on SQaLe
To isolate the effect of the data, we train Qwen3.5-2B with GRPO directly from the base checkpoint three times, once each on SQaLe, SynSQL-2.5M and BIRD train. We call the results MSQaLe (Qwen3.5-2B trained with GRPO on SQaLe), MSynSQL and MBIRD. Apart from the training data and the strength of a length curriculum, the base model, environment, tools, reward, steps, batch size and evaluation are identical. There are no distilled traces, no teacher and no test-time scaffolding.2
On the SQaLe test set (Figure 4), MSQaLe reaches 66.3%, which is 27.6 points above the untrained model and 9.7 points behind Qwen3.6-27B. Training on BIRD adds 15.3 points and training on SynSQL 12.0. On BIRD, MSQaLe reaches 52.3% against 54.7% for MBIRD, which was trained on BIRD itself. On EHRSQL the two are tied at 23.7%.
Execution accuracy (%) by training corpus
Figure 5 compares cost. All models run in the same agentic harness on 100 moderate SQaLe test questions. MSQaLe reaches 52% at 224 TFLOPs per question, and Qwen3.5-27B reaches 63% at 1,710 TFLOPs. Among general-purpose models in the 9–12B range, only Qwen3.5-9B (58%) is ahead of MSQaLe.
Accuracy against model size, 100 moderate SQaLe test questions
All 25 models as a table
How the model answers
MSQaLe never sees the schema. In each round it writes a short plan and then emits one or more tool calls. The environment runs every call, returns all the results in one message and reports how many rounds are left before the model has to answer. Figure 6 replays three of its episodes on full schemas of 106 to 239 tables, and all three end with a correct answer.3
All three episodes start the same way. The first round lists the tables, and the next one or two rounds inspect the few tables that match the question, with four to nine distinct calls in a round. The model then drafts a query and checks its output. The first two episodes end with submit_sql. The third keeps checking until the six rounds run out, and then the environment asks for the SQL as text. Most episodes end this way: in the 108-question evaluation these episodes come from, 84 of the 108 full-schema episodes reach the round limit.
MSQaLe also learned to batch its tool calls. The prompt asks for a single JSON object per reply, and we neither encouraged nor limited multiple calls; the environment simply runs every call it finds. During training, rounds per episode stay between 5.6 and 6.8 (Figure 7). Tool calls per episode stay near six for the first 800 steps, spike twice, and then climb steadily from step 954 to an average of 19.3 over steps 1,300 to 1,658. In the same evaluation, the untrained Qwen3.5-2B never issues more than one call per round on full schemas. MSQaLe issues several in 85% of its rounds, 5.5 calls per round of which 4.4 are distinct. The environment caps rounds, not calls, so batching lets the model inspect more of the database within the same budget. The other two GRPO runs picked up the habit less cleanly.
Rounds and tool calls per episode during training
Mean per reported step
What a small model learns
The three trained models differ only in their training data, so differences in how they behave come from the data. We look at three of them.
Schema size
Accuracy falls for every model as the schema grows, and MSQaLe leads at every size, from 72.7% with only the gold tables to 66.3% on the full schema. With only the gold tables present, all three models open every gold table in at least 95% of episodes. On the full schema, MSQaLe opens every gold table in 89.7% of episodes, against 85.0% for MBIRD and 83.3% for MSynSQL. BIRD and SynSQL training schemas have a median of 5 and 10 tables, so models trained on them never have to search.
Populated tables
Before submitting, MSQaLe has seen 88% of the string literals its answer depends on. MSynSQL has seen 55%, MBIRD 40% and the untrained model 35%. SynSQL’s tables hold a median of 2 rows, and SQaLe’s hold 69.
Schema size
Execution accuracy (%) at three schema sizes
String literals seen before submitting
Share of the string literals the answer depends on that the model saw in tool output (%)
Domains
On SQaLe test questions whose schemas lie outside BIRD’s domains, MSQaLe leads MBIRD by 9.9 points (46.4% against 36.5%). Inside BIRD’s domains the lead is 2.8 points (39.3% against 36.5%).4
Cite
If you use SQaLe, please cite this paper and the earlier workshop paper.
Notes
- On 158 BIRD dev questions (225 judge calls), 94.5% of the queries the judge accepts are aligned with the question. The judge agrees with human labels 86.2% of the time (κ = 0.68).↩
- The schema is withheld. The model explores the database through tools, the five used in answering plus
foreign_keysandjoin_path. The reward is tiered (correct result, then executes, then parses, then nothing), and partial credit never ranks a wrong query above a correct one.↩ - The episodes come from the full-schema condition of an earlier schema-size evaluation on 108 questions, not the 300-question evaluation behind Figures 4 and 8. Their questions use the evidence-supported phrasing, so each comes with a short evidence note.↩
- The domain split covers 108 schemas. That sample is too small to call the difference between the two gaps significant.↩