Methods
How Simparlia is built, how it is measured, and where it currently falls short. This is the technical companion to the research overview: it states the method in enough detail to be argued with, and publishes the numbers we hold — including the ones that are bad.
Method version 1.0 · 27 July 2026
01
Scope
Simparlia builds a model of each of the 650 sitting Members of Parliament from the public record, and uses it for two things: describing where an MP appears to stand, and predicting how they would vote on a division. This document covers the method behind both, and the measurements we have taken of it.
It is not the paper, which is in preparation. It is not a record of what shipped and when — that is the changelog. And it is not an accuracy report: no validated headline accuracy figure exists yet, so none is published here. What exists instead is a set of dated, committed measurements of the parts, and a protocol that says what would have to be true before an accuracy number could be published at all.
Every figure below carries the population it was computed on and the status of that population. Several rest on twenty MPs, or on a development split that deliberately excludes the held-out test set. Stated bare, they would mislead.
02
What we claim, and what we do not
Everything the model derives is presented as a model estimate, with the evidence behind it stated next to it rather than buried. Coverage is published, including where it is thin, and a surface that falls below the evidence floor is withheld rather than filled — an empty section with an explanation is more useful than a confident one resting on a single vote. Where a measurement has two readings we cannot separate, both are given and neither is asserted.
Not claimed
No accuracy figure. Independent validation is in progress and incomplete. Proof-of-concept accuracy figures quoted by stratum in earlier material have been withdrawn and are not restated here. They were computed before the holdout protocol in section 09 existed, on a sample that would not survive it, and against a whip signal that leaks the outcome.
No claim that source weighting improves predictions. The audit in section 04 shows that weighting changes which evidence dominates a judgement. Nothing yet tests whether the resulting judgements are better. That needs a validation run scoring predictions under weighted and unweighted evidence, and no such run has been made.
No claim that co-signature clusters measure agreement. They record that MPs put their names to the same motions. That is all they record.
No claim that a belief-versus-vote gap is hypocrisy. The gap has two causes the measurement cannot tell apart, and one of them is that our model is wrong.
03
The evidence base
Every source is public. Hansard speeches and interventions, select-committee evidence, bill sponsorships, every recorded division, Early Day Motion signatures, oral and written questions, registered financial interests, Electoral Commission donation records, the PublicWhip and TheyWorkForYou policy positions derived from voting records, self-published press releases, the MP’s own Bluesky posts, Wikipedia and Wikidata, and general web results. Nothing private or privileged is used, and nothing is bought.
The corpus is uneven in a way that matters. The cheap sources are near-universal — an encyclopedia entry, a career record, a donations row — and the expensive ones are not. An MP elected in 2024 may have a handful of speeches and no policy-vote history at all. Treating breadth of coverage as strength of evidence therefore rewards exactly the MPs about whom we know least of substance.
04
Weighing sources, not counting them
A synthesis that counts its sources gets the answer wrong. Three weak sources should not outvote one strong one, and the naive ordering is worse than that: it ranks a whipped division vote — the most official first-person record we hold — as our best evidence of what an MP believes. It is not. It is the best evidence of what the whip decided.
So each source family carries a weight that is the product of three factors. Authority: who published the record and how accountable they are for it. Voice: whose words these are and how considered they were. Attributability: whether the act reflects the MP’s own belief or an institutional constraint. Attributability is the load-bearing one. A whipped vote is maximally authoritative and minimally attributable; an Early Day Motion signature is coarser but freely given.
Source weights — selected anchors, on a 0–1 scale
- Hansard speech, committee evidence, bill sponsorship
- 1.000
- Basis — Tier A. The MP’s own composed words, on the official record, voluntarily.
- Early Day Motion signature, oral or written question
- 0.700
- Basis — Tier B. A deliberate act, freely chosen, yielding roughly one bit.
- Division vote, whipped or whip unknown
- 0.385
- Basis — Tier C. The default for every vote — the retriever cannot yet tell a free vote from a whipped one.
- Self-published press release
- 0.450
- Basis — Tier D. Deliberate and self-authored, but promotional and unaccountable. Ranking it above a whipped vote is the most contestable call in the table.
- Own Bluesky post
- 0.225
- Basis — Tier D. First-person, casual, often reactive. Reposts are excluded.
- Aggregated web or news result
- 0.063
- Basis — Tier E. Mixed, unvetted provenance — lowest confidence in what the source even is.
Recency is a separate multiplier rather than part of the weight, decaying to a floor rather than to zero: a 2015 speech is weaker evidence of a current belief, but it is not no evidence. Undated evidence is not penalised, because penalising it would confound “old” with “we failed to record a date”.
What re-scoring the existing evidence showed
An audit re-scored evidence the pipeline had already used, without re-running any synthesis, to find where a count and a weighted reading disagree.
Source-provenance audit, 27 July 2026
- Judgements whose most-numerous source family is not the most informative
- 24.6%
- Basis — 14 of 57 judgements across 20 MPs — the only profiles that record per-judgement provenance. Indicative of the mechanism, not a fleet rate.
- Judgements whose confidence band moves under weighting
- 26.3%
- Basis — 15 of 57. Same population.
- Judgements the old unweighted confidence score saturates at 100
- 73.7%
- Basis — 42 of 57. A defect in the instrument, not a property of the evidence: a cross-check that returns full marks three times in four is not a cross-check.
- MPs whose evidence-completeness tier changes when re-scored by weight
- 174 of 653
- Basis — Fleet-wide. Every one of the 174 moves downwards; rank correlation with the unweighted ordering is 0.909. Measures what evidence exists for an MP, not what any judgement used.
The disagreement runs one way and has one cause: division and policy-vote items are numerous and speeches are not. The retriever returns up to twenty vote items per dimension against a handful of speeches, so a count-based reading is led by the source with the weakest attributability while the MP’s own words sit underneath it.
Status of this model
The three factors and their ordering are the substantive claim. The numeric values are round numbers chosen to produce a defensible ordering — argued, not estimated from held-out accuracy. A validation run should replace them, and until one does they are placeholders.
The per-judgement source counts are what the model reported using, not a log of what the retriever returned. The retriever computes the true counts and does not yet persist them. Fixing that is the single change that would most improve the next audit.
05
Building a profile
Each MP is modelled in three layers. The first is core identity — biography, party, electoral history, constituency, ministerial role — taken from structured records. The second is belief and reasoning: ideological position on economic, social and sovereignty axes, plus stated positions by policy area, synthesised per dimension from evidence retrieved for that dimension across the whole corpus rather than from a sample. The third is contextual and injected at prediction time — the bill, its stage, the whip strength, the size of the government majority.
The model reasons as an analyst about an MP, not as a character playing one. Whip compliance is the default and rebellion has to be argued for from retrieved evidence, because that is the base rate in the Commons and because language models under-predict rebellion when asked to role-play. Output is structured rather than free text, so a prediction always carries its reasoning and cannot be a bare verdict.
Two profile generations, and why it constrains this page
653 profiles exist. 633 were built by the earlier single-pass path and record only a free-text evidence sentence per judgement — the source families behind those judgements were never recorded and cannot be recovered. 20 were built by the per-dimension path and record which families backed each judgement.
That is why the audit figures in section 04 rest on 20 MPs. It is the only population where the question can be asked at all. Re-running the audit after the fleet is rebuilt is a precondition of quoting any of those percentages as a fleet-wide rate.
06
Belief against vote
The same profiles support a second measurement, which appears in the product as conscience alignment. For every MP and policy area, the stance the model reads them as holding is placed on the same scale as the stance their voting record reveals, and the distance between the two is scored. Both sides run from −1 to +1; the gap runs from 0 to 2. The revealed side comes from PublicWhip division data; the stated side from a deterministic parse of the profile’s written positions, falling back to the ideology axis where no position parses.
Belief-versus-vote surface, method version 1.0, 27 July 2026
- MP × policy-area rows carrying both a stated and a revealed stance
- 5,142
- Basis — 613 MPs of the 650 with a belief-layer profile score at least one area.
- Mean absolute gap
- 0.446
- Basis — Median 0.442, on a 0–2 scale. Full fleet run.
- Rows where stated and revealed stances point in opposite directions
- 11.2%
- Basis — 575 of 5,142 rows. The record contradicting itself.
- …of those, where the revealed side is at least 80% whipped
- 57.9%
- Basis — 333 of 575. An association, not an explanation: the whip is the most common context for a contradiction, which is not the same as its cause.
- Mean gap on whipped votes, against free votes
- 0.43 / 0.65
- Basis — Divergence does not concentrate on whipped votes at the fleet mean — the opposite. Free-vote scores also rest on far fewer divisions, which inflates them mechanically.
- Correlation between an MP’s divergence and their rebellion rate
- −0.10
- Basis — Pearson, n = 606. The direction whip-following predicts, but small: rebellion rates are near zero for most MPs, so the signal is in the tails.
83% of scored rows fall back to the ideology axis rather than parsing a written position, and those rows diverge nearly twice as far as parsed ones — 0.48 against 0.27. That is a property of the fallback, not of the MPs, and it is the main reason some areas look more divergent than others.
How to read a gap
A non-zero gap has two causes that cannot be separated for any individual row. The MP voted with the whip against a view they hold. Or our model has mis-read the view. Both readings are on the table, and the product asserts neither.
They can be separated in aggregate, which is why the whipped share and the number of divisions travel with every score, and why areas backed by too few divisions are withheld rather than published thin. No row should ever be presented as “the model is wrong” or “the MP is a hypocrite” without that qualification.
07
Grouping without labels
Two kinds of faction claim appear on a profile, and they are kept apart because they are different kinds of assertion.
- Declared membership
- Read off a published roster, so a citation and an evidence date travel with every entry and an entry that arrives without one is dropped rather than rendered. Coverage is narrow by design: only groups that publish a membership list are curated. Groups that publish none — the European Research Group, the One Nation caucus — ship as empty stubs, because guessing at a roster would be inventing political affiliations for real people.
- Co-signature clusters
- Built by us, not read off anything. Every MP and every Early Day Motion form a bipartite graph; that graph is projected onto MPs, weighted by the number of motions two MPs both signed; community detection then runs within each party separately, so the communities are intra-party sub-groupings rather than the trivial party split. The seed is fixed and communities are re-ordered deterministically before identifiers are assigned, so repeated runs give identical output.
What a cluster is not
It is a record of behaviour: these MPs put their names to the same motions. It is not a measure of ideological similarity and not a claim that two MPs agree. Signing an Early Day Motion is cheap, public and low-stakes; MPs sign for constituency reasons, out of collegiality, because a campaign asked, or because no whip cared. Two MPs can share a hundred motions and disagree profoundly.
So the product publishes the raw overlap and nothing derived from it. There is no similarity score and no ranking, because inventing one would dress a signature tally as an ideological finding about a named living person.
No recovery rate against the declared rosters is published. Checking whether the clusters reconstruct a group whose membership is independently known is the obvious validation, and it is the right one. It has not been run as committed, reproducible code against the current rosters, so there is no number here to quote. An ad-hoc figure exists in our own working notes; it is not reproducible from anything in version control and we are not going to publish it.
08
Retrieval, measured
Retrieval is the component most able to fail quietly: a prediction built on the wrong evidence still reads as a confident prediction. So it has a committed baseline, dated and filed before any change it licenses, and the harness that produces it refuses to touch the held-out test set.
Relevance labels are structural rather than human-judged — debate membership, bill-title matches, sibling divisions under the same policy — which is a weaker standard than editorial judgement and is stated as such. Recall is macro-averaged across divisions, per retrieval stage.
Retrieval baseline, 27 July 2026 — recall at 5, 10 and 20, development split only
- MP-specific stage
- 0.417 / 0.550 / 0.592
- Basis — 41 of 73 development divisions carried structural labels; 32 carried none and were skipped. 621 labels in total.
- General-precedent stage
- 0.151 / 0.208 / 0.306
- Basis — Same divisions. This stage excludes the MP-stage hits by design, which depresses it somewhat — but it is the weaker half of the pipeline and we are not going to pretend otherwise.
Read plainly: on the MP-specific stage, a little over half of the structurally relevant evidence is in the top ten results. On the general stage, a fifth is. Both are development-set numbers on weak labels, and neither is a statement about how accurate a prediction built on that evidence turns out to be.
A canary that failed
Rebel-separability AUC — chance is 0.500
0.415
We tested whether an MP’s distance from their own party’s centroid in the embedding space separates rebels from loyalists. It does not. At 0.415 the score is below chance, which means the relationship runs weakly the other way: in the current whole-record embedding space, the MPs we label rebels sit marginally closer to their party centroid than loyalists do.
This is published because it is a real finding about the representation, and because any claim we later make about detecting rebellion has to be read against it. It is a canary on the embedding geometry, not a measure of prediction accuracy — nothing in the system predicts rebellion by centroid distance. But it removes an assumption we would otherwise have been entitled to make, and it means a probe or steering method built on this geometry would be building on sand.
Basis — Top-quartile rebellion definition; 156 of 629 MPs labelled rebels; 10 parties with at least three MPs. Retrieval baseline, 27 July 2026.
09
How this is validated
Synthetic-audience accuracy claims are worthless if the items being validated leaked into the data the model was conditioned on. The protocol below was adopted on 27 July 2026 and is binding on every simulation run and every accuracy claim we make. It exists so that our numbers can be checked, and so that the ones we have not earned yet are visibly absent rather than quietly assumed.
- The split is frozen
- 73 divisions are the development set and carry all iteration. 31 divisions are reserved for the paper — one pre-registered run per model configuration, with the full metric set declared in advance so there is no post-hoc metric selection. The evaluation harness refuses a test-set division by assertion, not by convention.
- The representation is frozen before a test run
- Embedding model, chunking and retrieval configuration are pinned before any test-set run, and cannot change between the run and the numbers reported for it.
- Conditioning is out of sample
- A prediction for a division must not condition on that division or its aftermath. The retrieval date cutoff is mandatory for every evaluation run; an ablation that disables it is labelled non-comparable rather than quietly compared.
- A known leak, stated rather than hidden
- Whip direction is currently inferred from party-majority behaviour, which indirectly reveals the outcome of the division being predicted. Every accuracy claim carries that caveat until the inference is replaced, and the final runs must state the method used.
- An accepted leak, also stated
- Profiles are built as of now over the whole corpus, so a profile can contain evidence postdating a validation division. Rebuilding a profile per division is not affordable. It is recorded as a threat to validity, to be quantified and revisited rather than waved away.
- Baselines are committed before the thing they gate
- Every measurement instrument files a dated baseline before the change it licenses is made, and the profile fleet is snapshotted before any regeneration so that before-and-after comparisons cite a fixed arm rather than a moving one.
None of this makes the system accurate. It makes the eventual accuracy number meaningful, which is a different and prior problem.
10
Known limits
These are the limits we know about. The list is not a disclaimer; each entry is a specific, actionable defect, and several are the next things to be fixed.
- Whip status is unavailable at retrieval time
- Every division is scored as though it were whipped, so free votes — assisted dying, abortion, the most attributable acts an MP performs — are under-weighted as evidence of belief. This is the highest-value outstanding fix.
- Per-judgement provenance covers 20 MPs
- Fleet-wide provenance figures measure what evidence exists for an MP, not what a given judgement drew on. The two must not be conflated, and section 04 keeps them apart.
- Policy positions carry no provenance at all
- Most of the judgement surfaces on a profile therefore cannot be weighted yet, whatever the tiering says.
- Era mismatch
- Stated positions describe an MP now; the voting record spans a whole career. Defectors, returning MPs and ministers whose brief changed produce large belief-versus-vote gaps that are neither model error nor current whip pressure.
- The sovereignty axis conflates two things
- It runs from internationalist to nationalist, which reads MPs of the Scottish, Welsh and Northern Irish parties as sovereigntist while their votes are pro-devolution and internationalist. That is a construct problem with the axis, not a bug in the pipeline, and it means those MPs’ sovereignty scores should be treated with more caution than the rest.
- Welfare scores for the 2019 intake are contested
- Two of our own ground-truth sources disagree with each other on the relevant policy, and in one case a categorical stance and a distance measure contradict outright. Divergence on welfare for those MPs should be treated with suspicion.
- Source weights are argued, not fitted
- Their ordering is the claim. Their magnitudes are placeholders awaiting a validation run that scores predictions under both readings.
- General-precedent retrieval is weak
- Recall@10 of 0.208 on the development split is the lowest measured number in the pipeline, and improving it is not optional.
- Rebellion is the known failure mode
- Language models pull toward the modal, expected answer, and rebellion is neither. Under-prediction of rebellion is the central risk to the whole approach, which is why the separability canary in section 08 is on this page rather than in a drawer.