Skip to content
Social Sense

Ask the data · evaluation

How often does the model get the SQL right?

16 questions about this database, 14 with gold SQL written by hand from the data and 2 that the data cannot answer. Run them on your model, with your key, and get execution accuracy with a confidence interval, the refusal rate, how often the validator had to step in, latency and token use.

Run

Benchmark your model

Add your own key to run the benchmark. Nothing is pre-scored or published here.

The 16 questions

Benchmark questions
idQuestionTests
q01How many geotagged tweets were matched to Victorian suburbs between February and July 2022?Sum over the 'all' topic only (topics overlap).aggregate
q02Ignoring the outlier filter, which Victorian SA2 has the highest median personal income, and what is it?All 457 Victorian SA2s, not only the 420 the IQR rule kept.ranking
q03How many Victorian local government areas did the team's IQR outlier rule keep?Uses the iqr_kept flag.aggregate
q04Which local government area recorded the most offences in 2019, and how many?Across all 79 LGAs, outliers included.ranking
q05What fraction (between 0 and 1) of crime-related tweets scored 1, the most negative score?A ratio over the Twitter crime histogram; integer division would return 0.ratio
q06What is the average sentiment score of all tweets in the suburb of Geelong?Needs the join from suburb name to SAL code.join
q07How many of the re-scored mastodon.social toots were declared as English?Language code 'en'.lookup
q08In the re-scored mastodon.social week, which hourly bucket (the hour_utc timestamp) had the most toots?One hourly bucket, as stored in hour_utc (not an hour of the day summed over the week).ranking
q09What Pearson correlation did the analysis find between median income and the sentiment of income tweets, across SA2s with at least one income tweet?Read the stored result rather than recomputing it.lookup
q10Which Greater Capital City area has the highest median income?From the capital-city summary table.ranking
q11List the five Victorian suburbs with the most crime-related tweets.Five names; order is not scored.join
q12How many Victorian SA2s have personal income data?457, before the outlier rule.aggregate
q13What is the mean sentiment score of income-related tweets across Australia?A weighted mean over the histogram.ratio
q14How many drug offences were recorded in the City of Greater Geelong?LGA names carry a type suffix.lookup
q15Which Twitter users posted the most crime-related tweets?No user data is stored: the right answer is a refusal.refusal
q16What was the average sentiment of tweets in Sydney during 2024?The tweets cover February-July 2022 only: the right answer is a refusal.refusal

Saved runs (this browser)

No runs yet. Completed runs are saved here for comparison.

Every model call of a run is also in the audit log.

Design

How the scoring works

Execution match. The model's SQL goes through the same validator and read-only runner as the live feature. Its result counts as correct when it contains the gold result's columns (in any order, extra columns allowed) and the same set of rows: integers must match exactly, other numbers within 0.006, text ignoring case. Row order is not scored, so “top five” is about which five.

Refusals. Two questions ask for things the database does not hold (individual users; 2024). Producing SQL for them counts as a miss even if the SQL runs. Declining an answerable question counts as wrong.

Uncertainty. Proportions carry Wilson 95% intervals; median latency a bootstrap interval. Fourteen scored questions give wide intervals by design: this is a smoke test that catches a broken prompt or a weak model, not a leaderboard. Repeat a run to see run-to-run variation, and compare two runs with the exact McNemar test on the questions where they disagree and a Tango score interval for the accuracy difference. A reply that is malformed or cut off counts as wrong; only key, quota, rate-limit, network or provider failures are left out.

No published scores. The site has no AI budget and does not report results it ran itself. Your runs stay in this browser (export them as CSV or JSON). The questions, gold SQL and pinned gold answers are in web/src/lib/sql/benchmark.ts and its test. More in methods.