# SQaLe: a large realistic dataset to empower small specialised text-to-SQL models Cornelius Wolff (1, 2), Daniel Gomm (1, 2), Madelon Hulsebos (2) (1) University of Amsterdam, (2) Centrum Wiskunde & Informatica Preprint, 2026 This is the plain-text version of the SQaLe project page at . The page draws its figures with JavaScript; here every figure is given as a table or a short description. - Questions and SQL: - Schemas and databases: - Model: - Earlier version: *SQaLe: A Large Text-to-SQL Corpus Grounded in Real Schemas*, - Index for LLMs: ## Abstract Frontier language models in agentic pipelines lead text-to-SQL benchmarks, at a high inference cost. Small specialised models would avoid that cost, but training them needs data that reflects the scale, semantics and structure of real databases. SQaLe is a semi-synthetic text-to-SQL dataset built on 9,259 real-world schemas from SchemaPile. Our generation pipeline extends each schema, fills it with synthetic rows, generates questions and answers them with an agent, and validates every stage by execution. The result is 1,408,056 natural-language questions paired with SQL, over schemas far larger than in any existing dataset. We train Qwen3.5-2B on SQaLe with GRPO alone, and its execution accuracy on the SQaLe test set rises from 38.7% to 66.3%. The model learns to explore the database: it finds the relevant tables in large schemas and reads the values its answer needs, which training on existing datasets elicits far less. ## Why SQaLe Small specialised models avoid the inference cost of frontier agentic pipelines. Training them needs data that looks like the databases they will be used on, and the common text-to-SQL training sets use small schemas. BIRD's training schemas have a median of 5 tables and SynSQL's a median of 10. A model trained on them never has to search for the tables a question needs. SQaLe starts from 9,259 real-world schemas from SchemaPile, extends each one with tables in its own style and fills it with synthetic rows. Its schemas have a median of 113 tables and 538 columns. Questions are generated from connected groups of up to 20 tables, so a model has to find the tables a question needs before it can answer. In the record in Figure 1, the question needs 3 of the schema's 87 tables. ## The dataset | | | |---|---| | question–SQL pairs | 1,408,056 | | real-world schemas | 9,259 | | distinct SQL queries | 176,761 | | median tables per schema | 113 | | median columns per schema | 538 | | synthetic rows | 108,708,694 | Each record pairs a natural-language question with a schema and a gold SQL query that has been run against the populated database. Every question exists in 8 phrasings, the original and seven rewrites in different styles, and all eight share one SQL query. The split into train and test is by schema, 95% / 5%. The dataset is published on Hugging Face in two parts that join on `schema_id`. [trl-lab/SQaLe-2-text-to-SQL-Queries](https://huggingface.co/datasets/trl-lab/SQaLe-2-text-to-SQL-Queries) holds the questions in all eight phrasings, the gold SQL and its result. [trl-lab/SQaLe-2-text-to-SQL-Schemas](https://huggingface.co/datasets/trl-lab/SQaLe-2-text-to-SQL-Schemas) holds the DDL and the generated rows of each database. ### Figure 1. One record from the test split `schema_013612` · 87 tables · 8,356 rows · moderate question · test split | Phrasing | Question | |---|---| | Verbose (original) | Identify the names of employees who have received feedback from the employee named "Tuttle Gledhill" and determine which department they belong to, in order to assess the reach of his feedback initiatives. | | Casual | Hey, can you pull up the names of everyone who got feedback from Tuttle Gledhill and tell me what departments they're in? Just need to see how far his feedback is spreading. | | Structured | 1. Identify all employees who received feedback from "Tuttle Gledhill". 2. Record their full names. 3. Determine the department for each employee. 4. Use this data to evaluate the reach of his feedback initiatives. | | Requirements list | Employees who received feedback from Tuttle Gledhill; their names; department assignments; feedback reach assessment. | | Short and high level | List the names and departments of all employees who received feedback from Tuttle Gledhill. | | Short and ambiguous | Track down who took feedback from Tuttle Gledhill and map out their teams. | | Spelling and grammar mistakes | Find the names of the peeps who got feedback from Tuttle Gledhill and figure out what deparment they are in, so we can see how his feedback inititives are doing. | | Evidence supported | Which employees received feedback from Tuttle Gledhill, and what are their departments?, evidence: Feedback records link a recipient to a sender. The sender is identified by the name 'Tuttle Gledhill'. | Gold SQL, the same for all 8 phrasings: ```sql SELECT DISTINCT e.first_name, e.last_name, d.d_name FROM feedback_session fs JOIN employee e ON fs.feedback_to = e.id JOIN department d ON e.department_id = d.id WHERE fs.feedback_by = ( SELECT id FROM employee WHERE first_name = 'Tuttle' AND last_name = 'Gledhill' ); ``` Result, 1 row: | first_name | last_name | d_name | |---|---|---| | Kilburn | Kempster | Identity Access Management | The schema has 87 tables. The gold SQL uses 3 of them: `feedback_session`, `employee` and `department`. SQaLe's databases hold 108,708,694 synthetic rows. Tables hold a median of 69 rows, 95.2% of tables are populated, and 95.1% of foreign-key cells are valid after repair. Every literal in a question is grounded in the data, so a model has to read values as well as table and column names. 23.4% of SQaLe queries are nested, against 7.7% in BIRD, and 40% of the queries that join chain multiple joins, against 26% in BIRD. 3.1% of queries touch five or more tables, up to 20. No BIRD or EHRSQL query touches more than four. ## How it is built The pipeline runs in five stages, and every stage is validated by execution. Value synthesis repairs rows that break key constraints, and answering sends failed or rejected queries back to the agent. 1. **Schema collection and extension.** SchemaPile provides real-world schemas from permissively licensed GitHub repositories. A tool-using agent annotates each of the 14,597 source repositories with a domain description. An LLM then extends each schema with tables that keep its naming conventions, normalisation level and foreign-key style. 2. **Value synthesis.** Tables are filled in foreign-key dependency order. For each table the LLM writes a Python function from its DDL, its original rows, the allowed foreign-key values and the domain description. Fact and junction tables get more rows and a skewed key distribution, and every table is checked for primary-key uniqueness and referential integrity. 3. **Question generation.** A connected subgraph of up to 20 tables is sampled along foreign keys, together with sample rows. The LLM writes questions that need every table in it, at three difficulty levels (simple, moderate, hard), with every literal grounded in the data. 4. **Style variation.** Each question is rewritten into the seven styles shown in Figure 1. All eight versions share the gold SQL. 5. **Agentic answering and judging.** An agent explores the live database with tools (`list_tables`, `describe_table`, `sample_rows`, `distinct_values`, `run_query`) and commits with `submit_sql`. Execution errors go back to the agent. An LLM judge checks the result against the question and sends rejected queries back for a rewrite.[^1] ## How it compares Figure 3 compares schema size across four datasets. The median SQaLe schema has 113 tables and 538 columns. The next largest medians are EHRSQL's 13.5 tables and 92 columns, measured on 2 schemas. SQaLe also has the most foreign-key relations, 1,196,078 across its 9,259 schemas. **Figure 3. Schema statistics.** | Dataset | Schemas | Median tables | Median columns | Foreign keys | Median rows / table | |---|--:|--:|--:|--:|--:| | BIRD | 80 | 5 | 39 | 526 | 3,738 | | EHRSQL | 2 | 13.5 | 92 | 34 | n/a | | SynSQL | 16,575 | 10 | 72 | 159,547 | 2 | | SQaLe | 9,259 | 113 | 538 | 1,196,078 | 69 | ## Training on SQaLe To isolate the effect of the data, we train Qwen3.5-2B with GRPO directly from the base checkpoint three times, once each on SQaLe, SynSQL-2.5M and BIRD train. We call the results M_SQaLe (Qwen3.5-2B trained with GRPO on SQaLe), M_SynSQL and M_BIRD. Apart from the training data and the strength of a length curriculum, the base model, environment, tools, reward, steps, batch size and evaluation are identical. There are no distilled traces, no teacher and no test-time scaffolding.[^2] On the SQaLe test set (Figure 4), M_SQaLe reaches 66.3%, which is 27.6 points above the untrained model and 9.7 points behind Qwen3.6-27B. Training on BIRD adds 15.3 points and training on SynSQL 12.0. On BIRD, M_SQaLe reaches 52.3% against 54.7% for M_BIRD, which was trained on BIRD itself. On EHRSQL the two are tied at 23.7%. **Figure 4. Execution accuracy (%) by training corpus**, on 300 SQaLe test questions, BIRD and EHRSQL. | Model | SQaLe test | BIRD | EHRSQL | |---|--:|--:|--:| | M_SQaLe (Qwen3.5-2B trained with GRPO on SQaLe) | **66.3** (+27.6) | 52.3 | **23.7** | | M_BIRD (Qwen3.5-2B trained with GRPO on BIRD train) | 54.0 (+15.3) | **54.7** | **23.7** | | M_SynSQL (Qwen3.5-2B trained with GRPO on SynSQL-2.5M) | 50.7 (+12.0) | 44.3 | 13.3 | | Qwen3.5-2B (untrained base model) | 38.7 | 19.3 | 8.2 | | Qwen3.6-27B (untrained, 27B parameters) | 76.0 | 69.3 | 55.0 | Figure 5 compares cost. All models run in the same agentic harness on 100 moderate SQaLe test questions. M_SQaLe reaches 52% at 224 TFLOPs per question, and Qwen3.5-27B reaches 63% at 1,710 TFLOPs. Among general-purpose models in the 9–12B range, only Qwen3.5-9B (58%) is ahead of M_SQaLe. **Figure 5. Accuracy against model size**, 100 moderate SQaLe test questions, all models in the same agentic harness. | Model | Parameters (B) | Accuracy (%) | TFLOPs per question | |---|--:|--:|--:| | M_SQaLe | 2 | 52 | 224 | | M_BIRD | 2 | 40 | 112 | | M_SynSQL | 2 | 33 | 155 | | Qwen3.5-2B | 2 | 24 | 168 | | Granite-4.1-3B | 3 | 21 | 414 | | Llama-3.2-3B | 3 | 10 | 535 | | Hunyuan-4B | 4 | 36 | 448 | | Qwen2.5-7B | 7 | 29 | 219 | | Olmo-3-7B-Instruct | 7 | 12 | 1,001 | | Olmo-3-7B-Think | 7 | 26 | 246 | | Llama-3.1-8B | 8 | 30 | 693 | | Qwen3-8B | 8 | 54 | 225 | | GLM-4-9B | 9 | 27 | 823 | | Qwen3.5-9B | 9 | 58 | 449 | | Gemma-3-12B | 12 | 39 | 717 | | Mistral-Nemo-12B | 12 | 14 | 1,239 | | Nemotron-Nano-12B | 12 | 45 | 504 | | Qwen3-14B | 14 | 54 | 369 | | gpt-oss-20b | 20 | 16 | 1,052 | | Mistral-Small-24B | 24 | 46 | 1,027 | | Qwen3.5-27B | 27 | 63 | 1,710 | | Qwen3-30B-A3B | 30 | 51 | 2,886 | | Gemma-4-31B | 31 | 60 | 1,867 | | Nemotron-Super-49B | 49 | 53 | 1,389 | | Llama-3.3-70B | 70 | 49 | 3,154 | ## How the model answers M_SQaLe never sees the schema. In each round it writes a short plan and then emits one or more tool calls. The environment runs every call, returns all the results in one message and reports how many rounds are left before the model has to answer. Figure 6 replays three of its episodes on full schemas of 106 to 239 tables, and all three end with a correct answer.[^3] **Figure 6. Three real episodes of M_SQaLe on full SQaLe schemas, with the schema withheld.** Identical calls within a round are listed once. **Shipping rates** · `schema_002386`, 239 tables · simple question · answered correctly - Question: What are the base and final shipping rates for PRIORITY_INTL for a 48.64 unit package? - Evidence: Base rate is the standard fee. Final rate applies weight-based adjustments. PRIORITY_INTL is an international shipping method. - Rounds: 1 Survey (`list_tables`) · 2 Inspect (`describe_table`, `run_query`) · 3 Query (`run_query`) · 4 Submit (`submit_sql`) - Final SQL: `SELECT sc.base_rate, sc.final_rate FROM shipping_rate_calculations sc WHERE sc.shipping_method_id = 'PRIORITY_INTL' AND sc.weight = 48.64`, which returns base_rate 23.9 and final_rate 71.51. - What to look for: The sampled rows of `shipping_rate_calculations` show that rates are stored under method codes such as PRIORITY_INTL, while `shipping_methods` uses names and SHM- identifiers. The model filters on the code it read. **Alert share** · `schema_018684`, 106 tables · hard question · answered correctly - Question: What percentage of all system alert log entries were triggered by user ID 5? - Evidence: Alert log entries are records of system notifications. User ID 5 is the specific administrator responsible for triggering the alerts. - Rounds: 1 Survey (`list_tables`, `run_query`) · 2 Inspect (`describe_table`, `run_query`, `sample_rows`) · 3 Draft and check (`run_query`, `describe_table`) · 4 Verify (`run_query`) · 5 Submit (`submit_sql`) - Final SQL: `SELECT (SELECT COUNT(*) FROM AlertLogs WHERE userId = 5) * 100.0 / (SELECT COUNT(*) FROM AlertLogs) AS percentage`, which returns 12.037037037037036. - What to look for: In round 3 the model runs the percentage query together with its two parts, 26 entries for user 5 and 216 in total, and checks one against the other. **Firmware devices** · `schema_007423`, 225 tables · moderate question · answered correctly - Question: List device names requiring software dependencies with 'firmware' in the name, along with their total logged diagnostic request counts. - Evidence: Device names are unique identifiers. Diagnostic requests are system logs. Firmware is a specific software component type. - Rounds: 1 Survey (`list_tables`, `run_query`) · 2 Inspect (`describe_table`) · 3 Read values (`describe_table`, `sample_rows`, `run_query`) · 4 Count (`describe_table`, `run_query`) · 5 Draft (`describe_table`, `run_query`) · 6 Check (`run_query`) · final answer written as text - Final SQL: ```sql SELECT d.name, (SELECT COUNT(*) FROM diagnostic_request_logs dr WHERE dr.device_id = d.uuid) as diagnostic_count FROM devices d JOIN device_software_dependencies ds ON d.uuid = ds.device_id WHERE ds.dependency_name LIKE '%firmware%' GROUP BY d.name ORDER BY diagnostic_count DESC ``` It returns Gateway-Remote Site-742 (2), Hub-Lab-252 (1), Sensor-Remote Site-776 (0) and Node-Lab-964 (0). - What to look for: The final query uses three of the 225 tables: it joins `devices` to `device_software_dependencies` and counts `diagnostic_request_logs` in a subquery. In round 3 a LIKE '%firmware%' query finds the dependency firmware-updater on four devices before the model writes any aggregate. Like most episodes, this one uses all six rounds and ends with the SQL written as text. All three episodes start the same way. The first round lists the tables, and the next one or two rounds inspect the few tables that match the question, with four to nine distinct calls in a round. The model then drafts a query and checks its output. The first two episodes end with `submit_sql`. The third keeps checking until the six rounds run out, and then the environment asks for the SQL as text. Most episodes end this way: in the 108-question evaluation these episodes come from, 84 of the 108 full-schema episodes reach the round limit. M_SQaLe also learned to batch its tool calls. The prompt asks for a single JSON object per reply, and we neither encouraged nor limited multiple calls; the environment simply runs every call it finds. During training, rounds per episode stay between 5.6 and 6.8 (Figure 7). Tool calls per episode stay near six for the first 800 steps, spike twice, and then climb steadily from step 954 to an average of 19.3 over steps 1,300 to 1,658. In the same evaluation, the untrained Qwen3.5-2B never issues more than one call per round on full schemas. M_SQaLe issues several in 85% of its rounds, 5.5 calls per round of which 4.4 are distinct. The environment caps rounds, not calls, so batching lets the model inspect more of the database within the same budget. The other two GRPO runs picked up the habit less cleanly. **Figure 7. Rounds and tool calls per episode during GRPO training on SQaLe** (run 26640986). The round limit never changes, so the growth in tool calls is growth in calls per round. The run trained for 1,800 steps; the reporting tool missed its updates after step 1,658, so the curves end there. Data: [episode shape (CSV)](https://trl-lab.github.io/sqale/assets/training/26640986-Episode-shape.csv). ## What a small model learns The three trained models differ only in their training data, so differences in how they behave come from the data. We look at three of them. ### Schema size Accuracy falls for every model as the schema grows, and M_SQaLe leads at every size, from 72.7% with only the gold tables to 66.3% on the full schema. With only the gold tables present, all three models open every gold table in at least 95% of episodes. On the full schema, M_SQaLe opens every gold table in 89.7% of episodes, against 85.0% for M_BIRD and 83.3% for M_SynSQL. BIRD and SynSQL training schemas have a median of 5 and 10 tables, so models trained on them never have to search. **Figure 8 (top). Execution accuracy (%) at three schema sizes**, with the share of episodes in which the model opened every gold table (%) in brackets. | Tables in the database | M_SQaLe | M_BIRD | M_SynSQL | |---|--:|--:|--:| | Gold tables only | **72.7** (95.3) | 64.7 (98.7) | 59.3 (98.3) | | +32 distractor tables | **68.0** (96.7) | 61.0 (90.7) | 52.3 (92.7) | | Full schema | **66.3** (89.7) | 54.0 (85.0) | 50.7 (83.3) | ### Populated tables Before submitting, M_SQaLe has seen 88% of the string literals its answer depends on. M_SynSQL has seen 55%, M_BIRD 40% and the untrained model 35%. SynSQL's tables hold a median of 2 rows, and SQaLe's hold 69. **Figure 8 (bottom). String literals seen before submitting:** the share of the string literals the answer depends on that the model saw in tool output. | Model | Literals seen (%) | |---|--:| | M_SQaLe | 88 | | M_SynSQL | 55 | | M_BIRD | 40 | | Qwen3.5-2B | 35 | ### Domains On SQaLe test questions whose schemas lie outside BIRD's domains, M_SQaLe leads M_BIRD by 9.9 points (46.4% against 36.5%). Inside BIRD's domains the lead is 2.8 points (39.3% against 36.5%).[^4] ## Cite If you use SQaLe, please cite this paper and the earlier workshop paper. ```bibtex @article{wolff2026sqale, title = {SQaLe: A Large Realistic Dataset to Empower Small Specialised Text-to-SQL Models}, author = {Wolff, Cornelius and Gomm, Daniel and Hulsebos, Madelon}, journal = {arXiv preprint}, year = {2026} } @article{wolff2025sqale, title = {SQaLe: A Large Text-to-SQL Corpus Grounded in Real Schemas}, author = {Wolff, Cornelius and Gomm, Daniel and Hulsebos, Madelon}, journal = {arXiv preprint arXiv:2602.22223}, note = {AI for Tabular Data workshop at EurIPS 2025}, year = {2025} } ``` ## Notes [^1]: On 158 BIRD dev questions (225 judge calls), 94.5% of the queries the judge accepts are aligned with the question. The judge agrees with human labels 86.2% of the time (κ = 0.68). [^2]: The schema is withheld. The model explores the database through tools, the five used in answering plus `foreign_keys` and `join_path`. The reward is tiered (correct result, then executes, then parses, then nothing), and partial credit never ranks a wrong query above a correct one. [^3]: The episodes come from the full-schema condition of an earlier schema-size evaluation on 108 questions, not the 300-question evaluation behind Figures 4 and 8. Their questions use the evidence-supported phrasing, so each comes with a short evidence note. [^4]: The domain split covers 108 schemas. That sample is too small to call the difference between the two gaps significant. Contact: {cornelius.wolff, daniel.gomm, madelon.hulsebos}@cwi.nl. Website content MIT licensed. --- # Using SQaLe: data, models and code This is the reference part of the plain-text SQaLe documentation. The project page itself is at . ## The data SQaLe is published as two Hugging Face datasets that join on `schema_id`. | Dataset | Contents | Rows | |---|---|--:| | [trl-lab/SQaLe-2-text-to-SQL-Queries](https://huggingface.co/datasets/trl-lab/SQaLe-2-text-to-SQL-Queries) | questions in eight phrasings, gold SQL, difficulty, the gold query's result | 177,377 | | [trl-lab/SQaLe-2-text-to-SQL-Schemas](https://huggingface.co/datasets/trl-lab/SQaLe-2-text-to-SQL-Schemas) | the DDL and generated table rows of each database | 9,259 | Both are split by schema into `train` (8,836 databases, 169,277 question records) and `test` (423 databases, 8,100 question records), so no test database appears in training. The data is in parquet files, MIT licensed. ### Fields of trl-lab/SQaLe-2-text-to-SQL-Queries | Column | Type | Content | |---|---|---| | `question_id` | string | unique id of the question record | | `schema_id` | string | join key into trl-lab/SQaLe-2-text-to-SQL-Schemas | | `sql` | string | the gold SQL query (SQLite) | | `difficulty` | string | `simple`, `moderate` or `hard`, the target level the question was generated for | | `questions` | struct | the question in eight phrasings (below) | | `relevant_tables` | string | JSON list of the tables in the subschema the question was generated from; the gold SQL uses a subset of them | | `number_of_relevant_tables` | int | length of `relevant_tables` (1 to 20, median 5) | | `execution_result` | string | JSON list of up to the first 50 result rows of the gold SQL on the populated database | The eight phrasings in `questions`: `verbose` (the original, fully specified question), `evidence_supported` (a compact question followed by `, evidence: ` and a short note with the outside knowledge needed to map it onto the data), `structured` (bulleted or numbered requirements), `requirements_list` (only fragments naming what is wanted), `short_ambiguous` (a short version that leaves part of the specification implicit), `short_high_level` (a short paraphrase of the top-level intent), `casual` (an informal restatement) and `spelling_grammar_mistakes` (the question with typing and grammar errors). ### Fields of trl-lab/SQaLe-2-text-to-SQL-Schemas | Column | Type | Content | |---|---|---| | `schema_id` | string | join key into trl-lab/SQaLe-2-text-to-SQL-Queries | | `full_schema` | string | the DDL: every `CREATE TABLE` statement of the database in one string | | `schema_content` | string | JSON object mapping each populated table to a list of row objects (`{"column": value, ...}`); unpopulated tables are absent | | `number_of_tables` | int | number of `CREATE TABLE` statements in `full_schema` | ### Loading the data The SQaLe Python library ([PyPI](https://pypi.org/project/SQaLe/), [source](https://github.com/trl-lab/SQaLe-Library)) reads both datasets and writes the databases as SQLite files: ```bash pip install "SQaLe>=0.2" sqale-extract --split test --output ./dbs ``` ```python import sqlite3 from sqale import deserialize_sqale, load_questions questions = load_questions(split="test", limit=100) databases = deserialize_sqale(split="test", output_dir="./dbs", schema_ids={q["schema_id"] for q in questions}) db_path = {d["schema_id"]: d["db_path"] for d in databases} q = questions[0] conn = sqlite3.connect(db_path[q["schema_id"]]) print(q["questions"]["verbose"]) print(conn.execute(q["sql"]).fetchmany(5)) ``` `load_questions` parses the JSON columns and replaces placeholder phrasings with `None`. The parquet files also load directly with `load_dataset("trl-lab/SQaLe-2-text-to-SQL-Queries")` from the `datasets` library. ### Known issues - 87,365 of the 1,408,056 phrasings are placeholders left by the style-variation step, such as `...`, `` or a bare difficulty label, and 1,752 records have no usable `verbose` question. Filtering them leaves 1,320,691 phrasings. - Table rows are generated by an LLM, not collected. Tables hold a median of 69 rows, and 4.8% of tables are empty. - Gold queries were accepted by an LLM judge whose accepted queries were aligned with their question in 94.5% of the validation calls. - The SQL targets SQLite, and all questions are in English. ## The models Three Qwen3.5-2B models trained with GRPO from the base checkpoint, one per training corpus, with everything else identical: | Model | Training corpus | |---|---| | [trl-lab/qwen3.5-2b-grpo-sqale](https://huggingface.co/trl-lab/qwen3.5-2b-grpo-sqale) (M_SQaLe) | SQaLe | | [trl-lab/qwen3.5-2b-grpo-bird](https://huggingface.co/trl-lab/qwen3.5-2b-grpo-bird) (M_BIRD) | BIRD train | | [trl-lab/qwen3.5-2b-grpo-synsql](https://huggingface.co/trl-lab/qwen3.5-2b-grpo-synsql) (M_SynSQL) | SynSQL-2.5M | The models are agents, not single-shot SQL generators. Each model repository includes `sqale_agent.py`, a reference loop that reproduces the training environment and opens the database read-only: ```bash pip install vllm huggingface_hub hf download trl-lab/qwen3.5-2b-grpo-sqale sqale_agent.py --local-dir . python sqale_agent.py --model trl-lab/qwen3.5-2b-grpo-sqale --db path/to/database.sqlite \ --question "Which supplier shipped the most back-ordered items last quarter?" ``` The protocol, for anyone writing their own loop: - The schema is never in the prompt. The system prompt holds the task, the tool list and an execution-plan scaffold; use the exact strings `SYSTEM_PROMPT` and `USER_TEMPLATE` from `sqale_agent.py`. - After a `` block, each assistant turn holds one or more bare JSON tool calls, `{"tool": "", "args": {...}}`. - Tools: `list_tables`, `describe_table(table_name)`, `foreign_keys(table_name)`, `join_path(table_a, table_b)`, `sample_rows(table_name, limit)`, `distinct_values(table_name, column_name, limit)`, `run_query(sql)` and `submit_sql(sql)`, which ends the episode. - Tool results come back as the next user turn, followed by the remaining budget, for example `[4 tool turns left before you must answer]`. - After six tool turns, or four failed `run_query` calls in a row, the model gets one final turn to answer. An empty or broken final answer falls back to the last query that ran without error. - Earlier `` blocks stay in context, so append completions to the raw prompt rather than rebuilding it from a message list. - Sampling: temperature 0.6, top-p 0.95, top-k 20; stop on `<|im_end|>` and `<|endoftext|>`. An episode is budgeted at 12,288 tokens.