Scoring Methodology

How policies are read, scored, and classified — and where the method is still imperfect.

What the scores represent

Each policy is scored on five dimensions, each on a scale of −10 to +10:

DimensionWhat it measures
SocialEffect on people’s welfare, rights, health, education, housing
EnvironmentalEffect on climate, nature, emissions, energy, pollution
EconomicEffect on the material wellbeing of the general public: prices paid, incomes earned, employment, sustainability of public finances
Human RightsEffect on civil liberties, freedoms, and legal rights
GovernanceChanges to who holds power, how it is checked, or what the public can see and contest

Scores reflect magnitude of real-world impact, not whether a policy is ideologically “good” or “bad”. A law that dramatically cuts a public programme scores high (negatively) on social. A law that transforms environmental regulation scores high (positively or negatively) on environment.

null means the dimension is not applicable — not “uncertain.” A fisheries quota has no governance dimension; an anti-corruption law has no environmental dimension.

Scoring versions

Where the site is now: The rescore to v4 is well underway — recent activity (what you see on the homepage by default) is almost entirely v4 already, since new records are scored directly under v4 and the rescore works backward through history. Older records further back in time are still being converted from v3 to v4 in the background, record by record. The footer of the main page shows current progress. This page explains what changed and why.

v3 (being replaced)

The v3 corpus used a single prompt but was scored across roughly a dozen different models as free-tier quotas allowed. DeepSeek scored approximately 79% of records; the remainder were split across Gemma 4 31B, Gemini Flash-Lite, Gemini 3.1 Flash-Lite, Groq, and several Qwen variants. Those models calibrate differently — most visibly on when to return null rather than a number. On human rights scores alone, the null rate ranged from 0% (Gemma never returns null) to 100% (Qwen always returns null for this dimension) to 27–38% (DeepSeek and Gemini).

Individual v3 scores are usable as rough indications of a law’s impact. V3 aggregates and trend comparisons are not reliable, because they partly reflect which provider was available in a given scoring period rather than real-world changes in policy direction. This is also why the country comparison table is suppressed — country rankings would partly measure which country’s laws happened to be scored by which model, not how those laws actually compare.

What the study found: An inter-provider agreement study (September 2026) measured how often two independent models agreed on the direction of each score on a 300-record sample. Economic scores disagreed 31.4% of the time, against a same-provider noise floor of 6.7% — meaning roughly 25 percentage points of that disagreement was genuine divergence driven by prompt ambiguity, not model randomness. On the same sample, switching from the v3 to the v4 prompt caused 17.8% of economic scores to flip sign. This confirms v3 and v4 economic scores are not comparable, and that a full rescore was necessary rather than a partial patch.

v4 (current)

Economic score: scored from the perspective of ordinary households — who bears the cost, who receives the benefit. A law that benefits a narrow industry while spreading the cost across the public scores negative regardless of aggregate GDP effect. A minimum wage increase scores positive even if it raises employer costs. Decision procedure follows seven explicit rules covering distribution of costs and benefits, with ten worked cases covering the specific cases that produced model disagreement under v3.

Governance score: scored only if the law changes (a) who holds decision-making power, (b) how that power is checked, reviewed, or overturned, or (c) what the public can see, access, or contest. The default is null. Administrative procedures, budget allocations, departmental reorganisations, and routine appointments do not meet this threshold. Eleven worked cases define the boundary.

Null: models are instructed that null means “not applicable” and that it is normal for most dimensions of most laws, but never for all five at once. Under v3, models defaulted to scoring every dimension with a low number rather than returning null, which produced spurious signal in aggregate charts.

Never all five null (since 29 September 2026): a law we score gets at least one number, with 0 meaning a real but negligible effect. A document where no dimension applies is shown as not scored, with the reason “no measurable effect”. Before this rule, about 39% of scored laws came back with no number at all, and some real rules were missed. Those laws are being scored again.

Helped/hurt classification thresholds:

Provider and model

All v4 scores use Gemini Flash-Lite (gemini-flash-lite-latest) at temperature=0 with structured output. The rescore changes both the prompt and the effective calibration: the v3 corpus used roughly a dozen models with inconsistent null behaviour; the v4 corpus uses one. Single-provider scoring is what makes trend data and cross-law comparisons meaningful, because the null calibration and score magnitude are consistent across the entire corpus.

Country rankings

The country ranking table is not currently published on the public site. The inter-provider study found a Spearman ρ of 0.411 between the v3 mixed-provider corpus and a second model — but that figure reflects the v3 corpus’s multi-provider noise, not inherent model disagreement. Whether a single-provider v4 corpus can support a reliable country comparison table is an open question to revisit once the full rescore completes. For now the table remains suppressed.

Source quality and evidence levels

Not all policies arrive with the same amount of text. Records are classified internally:

LevelCriteriaNotes
full_textraw text > 2,000 charsSubstantial text available; most reliable
abstract201–2,000 charsSummary or abstract available
title_onlyraw text ≤ 200 charsScored from title alone; least reliable; excluded from all aggregate charts

EUR-Lex records typically arrive with title and CELEX number only (title_only). UK legislation.gov.uk records arrive with up to 4,000 chars of full text. This inconsistency is a known limitation and is a separate source of variance from model or prompt choice.

Decisions we show but never score

Some parliamentary decisions are about one named person rather than a policy — for example, lifting or keeping a member of parliament’s immunity so a court case can go ahead. Their text is a short template with a name in it, so an AI score could only reflect guesses about the person, not the decision. We tested this in September 2026: identical resolutions received opposite scores. We still show these decisions, because they are real acts of parliament, but we never score them, and they are left out of all charts and statistics.

Other official papers aren’t policies either, so we don’t score them: appointments and dismissals of named people, honours and decorations, corrections of earlier texts, re-publications, notices that a court case was filed, routine technical notices and gazette indexes. They are kept out of the feed; the archive shows them, with the reason, when you tick “include routine paperwork”. The AI can also decline to score a text itself, when the text is not a policy or doesn’t say enough to judge; it then writes only a short summary.

Laws are also never scored from their title alone: a law waits unscored until we have its text.

What is not published

See also: data coverage — where the records come from, how far back, and known gaps.