methodology · the long version

Methodology

Tayyar locates MENA political parties, politicians, and activists in ideological space across 16 dimensions. Every claim about every entity links back to a source you can check; every score follows from a rubric you can read. It treats the region on its own terms — the rubric is anchored to MENA reference points, not Western defaults. This page is the complete account: how the dataset is built, what's verified, where it falls short, and what's planned. It's deliberately exhaustive — use the contents to jump.

01What Tayyar is

Tayyar is a research dataset that places the actors of Middle Eastern and North African politics in a shared, sourced coordinate system. Where a standard left–right axis flattens the region into a caricature, Tayyar scores each actor on 16 dimensions chosen to surface the cleavages that actually move MENA politics — state and religion, the regime/resistance question, the Palestinian cause, sectarian power-sharing, Iran posture — and pairs every position with the evidence and the reasoning behind it.

Scope is tiered. A Tier-1 set (Egypt, Israel, Palestine, Lebanon) is the deepest; Tier-2 fills out the remaining states, 20 in all, from Mauritania to Oman. The dataset is in a pilot phase: today's positions are first-pass estimates, in the process of being upgraded to document-grounded scores synthesised from a multi-rater panel. Nothing here is presented as settled truth — it's presented as reproducible, sourced, and honestly caveated.

Two questions run through everything and are kept strictly apart. Verification asks whether the facts about an entity are correct (its founding year, its leader, its legal status). Reliability asks whether independent raters place it in the same position. A perfectly verified party can still sit on a contested axis, and the page never lets one stand in for the other.

02The entity model

The dataset is built from four entity types, each with per-field citations and a status taxonomy.

  • Countries (20) — the geographic frame; each carries its own page and coverage rollup.
  • Parties (98) — one row per party. Joint lists and electoral alliances are modelled as their own entity with member-party edges — Hadash-Ta'al links to Hadash and Ta'al; The Democrats links to Labor and Meretz. A joint list that spans ideologically distinct members carries no single compass position (averaging would mislead); a merger of close parties keeps one. 4 such member relationships are recorded.
  • Politicians (274) — heads of state and government, party leaders, ministers, sitting MPs and MKs, opposition figures, activists, and dissidents. A politician links to a current party and, today, inherits that party's position unless scored in their own right; where an individual is document-scored, the page renders their own profile instead.
  • Axes (16) — the dimensions themselves are first-class records, each with a definition, pole labels, scoring conventions, and anchor examples, so the rubric is data, not prose buried in a script (§3).

Historical and dissolved parties are included and scored, not omitted. Each is marked inactive and its position is stamped to its era (the dataset is temporal), so it stays out of the "current landscape" views by default but is available for cross-era comparison — Ben-Gurion's Mapai against today's Likud. The axes are largely time-invariant and the anchors are themselves historical (Nasserist Egypt), which keeps such comparisons honest as long as the era is explicit.

03The 16 axes

The standard two-axis economic/social compass collapses the distinctions that define the region. Tayyar keeps those two as the default 2-D plot and adds 14 issue axes that populate the per-party spider and have their own deep-dive pages. Each axis runs −10 to +10, oriented by its own poles — 0 means genuinely mixed or not engaged, never "centrist by default." Each carries a definition, MENA-anchored scoring conventions (Nasserist Egypt at the statist −10; Saudi Vision 2030 at the market +10), and anchor examples — all stored in the database and fed verbatim to the analyzer, so the tool scores exactly the dimensions the dataset does.

The four governance-adjacent axes were sharpened after the reliability pilot (§11) to remove ambiguity: liberal democracy means commitment to liberal checks, not merely holding elections; west-alignment is strategic, not cultural; regional stance is the resistance/normalization axis, not sympathy; social is moral order, not nationalism.

Compass axes (2) — the default plot

Issue axes (14) — the spider profile

Applicability — not every axis applies to every actor

An axis can be meaningless for an actor: pan-Arab vs particularist doesn't apply to a Jewish-Israeli party, monarchical legitimacy doesn't apply in a republic, normalization with Israel is an Arab-actor question. Scoring those cells 0 would be a false reading, so each axis carries a scope and the scoring and rendering skip the cells it doesn't reach: universal (all actors), arab-actors, israel, monarchies, confessional states. This is why a spider may legitimately leave a spoke blank rather than plant it at the centre.

Context axes — defined, activating with the next scoring round

The axis set is being widened with context-specific dimensions, locked into the frame now but held inactive until the fresh scoring covers them (so they don't render as empty spokes in the meantime):

  • Normalization with Israel arab-actors Anti-normalization ← → Pro-normalization Stance of Arab actors toward diplomatic and economic normalization with Israel.
  • Territories & sovereignty israel Annexation / Greater Israel ← → Withdrawal / two-state Israeli actors' position on the West Bank and Gaza: annexation and settlement vs territorial compromise and a Palestinian state.
  • Jewish–Arab equality israel Ethnonational Jewish state ← → Full civic equality Within Israel: a Jewish ethnonational state vs full civic equality for Arab/Palestinian citizens.
  • Monarchical legitimacy monarchies Republican / anti-dynastic ← → Monarchist / dynastic In monarchies (Jordan, Morocco, the Gulf): acceptance of dynastic royal legitimacy vs republican or anti-monarchical challenge.

04The lens system

A single number hides as much as it shows. A position can be read three ways, and the spread between them is itself a finding:

Declared

What the actor says — platforms, manifestos, charters, official rhetoric. Sourced from the actor's own documents.

Behavioral

What the actor does — votes, coalition choices, the governing record. The test of declared positions.

Perceived

How the actor is seen — third-party characterisations, where they materially differ from the self-image.

Composite

The published synthesis. What the compass and map render unless you toggle a lens.

The operating rule, fixed for the scaled rounds: the declared lens is the spine — every scored cell carries it, grounded in the actor's own documents, because a stated position is what can be cited and reproduced. The behavioral lens is scored where it bites: for actors who hold or have held power, on the conduct axes (democracy, civil liberties, press freedom), where a promise is cheap and the record is the real test. There a charter's pledge and a governing record can point opposite ways, so the composite takes the behavioral reading; everywhere else it equals the declared. Actors whose character changed across eras are scored as distinct entities — a 1947 founding doctrine is not its later regime — so no single cell is asked to hold a contradiction.

Most cells carry only the declared reading today; the behavioral split fills in for power-holders as the scaled rounds land. The compass already carries a lens toggle, and 20 parties show a measured declared-vs-behavioral gap on at least one axis — that seam, not either number alone, is the most interesting thing on the page.

05How a score is made — the rater panel and synthesis

Every published position is the synthesis of a rater panel, not one person's call. The panel is built from four kinds of rater — the author-baseline prior and the nine-model language panel active today, with blind human anchors planned for r3 — and each published score is the median of whatever ratings actually exist for that (entity, axis) cell, carrying a confidence, an agreement figure, and a resolution status.

RaterKindCoverageRole
author-baselinepriorevery cellthe v0 hand-coding that seeds the panel; retired from the median once a human vote lands, kept as a labelled prior. Today's only non-model signal.
Nine language modelsLLMr2 panelClaude, Gemini, GPT, Gemma, Grok (US) and Kimi, DeepSeek, Qwen, MiniMax (China) — each scores blind; cross-lab agreement reduces single-model idiosyncrasy, but shared training data means it cannot rule out a common prior (see §11)
Grounded pilotLLMreferenceone model scored directly from cited documents; kept as a document-verification reference, not a panel vote
Tarek, second raterhumanplanned · r3two blind human anchors → the human-to-human ceiling. Not yet collected; the panel today is models plus the author-baseline prior.

They are collected in that order to preserve blinding — grounded first, then each model blind, with the human anchors planned as the final layer, blind to both — and stored in one panel table (tayyar_ratings). Synthesis is two-stage: first collapse each rater's multiple rounds to their own median, then take the panel median across distinct raters, so a model that rated three times still counts once. Each cell then gets:

  • score — the panel median.
  • agreement — 1 − (spread ÷ 20) when ≥2 raters exist.
  • statussettled when ≥2 raters land within tolerance (MAD ≤ 1.5) and at least one is document-grounded; contested when they diverge sharply (MAD ≥ 3 — the chart then shows a range, not a point); provisional otherwise, which is most cells today because the panel is still filling in.

Live panel status across 1108 published cells, from 16287 ratings by 14 raters over 8 rounds: 795 settled 255 provisional 58 contested · 1004 cells already have ≥2 independent raters — browse the full panel, every rater's score per cell. Crucially, the more documents we seed, the more cells get a grounded rating, so coverage measurably raises confidence — scoring quality is tied directly to the source corpus.

From cells to field-level claims — the inclusion rule

A cell's status governs that cell. Field-level numbers are a second layer: the "most contested" axis rankings, the headline cross-axis correlations, and the family and country aggregates as they come online. A sparse axis can quietly poison them — with only a handful of entities scored, one outlier swings the mean, a high standard deviation reads as "contested" when it is merely thin, and a spurious correlation can top the table.

So an axis earns a place in field-level aggregates only when its coverage clears a floor (currently 30 scored entities). Coverage is the structural gate because it is what makes an average mean anything. Cross-rater dispersion rides alongside as a distinct signal — a well-covered axis can still be one the panel splits on, and is flagged rather than demoted, while a sparse axis can be tightly agreed where it does apply — but coverage decides eligibility. Axes below the floor are reported in full, on their own, never folded into a headline: the sparsity marks where coverage is thin, not where the region agrees. The findings page shows the split live.

06The multi-round rater protocol

The panel fills in over a defined sequence. Each LLM rates in fresh, blind chats, and the same axis definitions ship inside the rater-console template so every rater — model or human — sees identical instructions. The rounds are labelled in the data (8 present so far) and synthesised as in §5.

  1. r1 — the "before." Gemini + Grok, without axis definitions, kept deliberately so the later rounds can measure whether definitions improved agreement. author-baseline supplies the v0 prior.
  2. r2 — definitions in. Claude + Gemini + Grok + Kimi + DeepSeek in fresh chats, all working from the sharpened axis definitions. The live round.
  3. r3 — the human anchors (planned). Two blind human raters (Tarek, then a second), over a stratified sample, giving the human-to-human ceiling and the panel's first non-model votes.
  4. r4 — blind LLM-as-judge (planned). A separate, held-out frontier model over three cycles, fed the rater outputs with anonymised, shuffled labels (A/B/C/D) so it cannot favour any single model; it scores the divergence and which placement is most defensible, and the median of its cycles enters the panel.
  5. r5 — synthesis & check (grounded pilot run; full pass planned). The full matrix is synthesised and verified against the seed documents specifically for declared-vs-behavioral divergence.

Because each model rates across multiple rounds, the protocol also yields test–retest self-consistency — a model that flip-flops between its own rounds is flagged. Ratings are added through one channel: the grounded pilot's ratings are written directly; every other rater goes through an internal Rater Console, a single blind-CSV round-trip — download a template with no other scores in it, fill it (humans in a spreadsheet, the other models via a copy-paste prompt block), re-upload, and the panel re-synthesises. This is the part of the work currently paused; the substrate, console, and synthesis are built and the rounds resume on demand.

07The source-document pipeline

Scores are only as good as the texts behind them. A separate hunter agent gathers verified primary material and stages it in a review queue; Tarek approves; a seeder routes approved rows into the production tables. 426 primary texts are seeded so far, across three original languages. The hunter gathers six kinds of material:

documentsspeecheseventsbillsbiossocial

Social is the newest kind — tweets and posts by parties and politicians, which are often the most candid evidence of a position or strategy. It carries one extra rule the others don't: an archival snapshot (an archive.org capture or screenshot, with the capture date and the verified account handle), because posts get deleted — verification has to assume the original will vanish.

The bar for every item:

  • Verbatim where it's a text — the excerpt is copied from the source, not paraphrased; for events/bills/bios every claim is backed by the cited source.
  • A durable, credible host — government and parliament archives, party sites, MFA / royal-diwan pages, the UN, court records, reputable outlets. Link-rot hosts are rejected.
  • Original delivery language recorded. A speech delivered in Arabic is tagged ar even when only an English translation is stored — the language field is what it was delivered in, never defaulted. English-language sources are first-class.
  • A source-authority tier — 1 (first-party / official / canonical), 2 (reputable outlet / mirror), 3 (aggregator citing a primary).
  • A known date, approximate years flagged.

Every document is also evidence. Before a row is accepted the hunter flags any contradiction between the source and the existing Tayyar record (verification flags) and promotes individually striking, axis-engaging sentences to the quote corpus (candidate quotes). The hunter never edits the dataset — both feed a human review pass that decides what changes.

08Document-grounded scoring

The scoring pipeline reads a seeded document and scores it across every axis it engages, citing the verbatim spans that drove each score, then writes one declared-lens analysis row per axis and recomputes the entity's aggregate. This is what turns the author-baseline priors into grounded scores that carry per-passage provenance — a researcher can follow any score back to the sentence behind it.

Where both a hand-coded composite and a document-grounded score exist for the same (party, axis), the pair is an inter-method check: mean absolute difference and Pearson r per axis are computed live and published on findings. Today the corpus is just large enough to begin populating that table; as parties accumulate grounded scores the agreement numbers appear automatically, and that section will lead the paper's argument: which dimensions of MENA politics resist coding, and why.

09The paste-and-position analyzer

/tayyar/analyze runs the same scoring engine on any text you paste — a manifesto excerpt, a speech, an op-ed. It detects the language, scores every axis the text engages (the axis catalogue is pulled live from the database, so it scores exactly the dataset's axes), and returns, per axis, a score, a confidence, one-to-two sentences of reasoning, and the verbatim spans that drove it. On top of that it gives a plain-language synthesis, a spider profile, the cited phrases highlighted in place by polarity, and the parties or politicians the text most reads like (nearest by RMS distance across the engaged axes, with the option to overlay one on the radar).

It is pinned to a model and a prompt version (analyze-v0.3.0) so any analysis reproduces, and the model is forced to return a structured schema — no free-text parsing, the result shape is deterministic. Anonymous users get 3 scores a day; signed-in users get 30 and can save an analysis to a citable permalink. The analyzer is the dataset's methodology made interactive: the same rubric, the same axes, the same insistence on a quotable span behind every claim.

10Verification & provenance

For every party, politician, and country, individual facts are tied back to where they came from. Each entity carries a list of citations — one per specific claim, not one per row — each saying "this leader name comes from this URL, accessed on this date." So when you read that a given figure is currently imprisoned, you can follow the citation to confirm it.

Verification runs in two passes. First-pass: review against widely-available reference and official party material, citations attached, entity marked reviewed — the dataset is 51% of the way through. Second-pass: a careful human check against the original primary sources, planned before any formal publication and backed by the 426-document corpus. Status climbs unverifiedtarek-verified (first-pass) → externally-verified.

The compass scores remain hand-coded priors marked unverified until they pass the panel and grounding pipeline. Verification validates the facts about an entity; it does not, on its own, validate a placement. The two are never conflated.

11Inter-rater reliability

Before the panel design was trusted it was stress-tested on a stratified sample re-coded across three rounds by two independent second-pass coders, blind to the baseline. Per-axis quadratic-weighted Cohen's κ, Pearson r, and mean |Δ| are live on findings §7.

What it found. Strong agreement overall, but the pattern is the signal. The conflict-defining axes — Palestinian question, west-alignment — reproduce almost perfectly. The governance axes — democracy and economic — carry several times the disagreement, because scoring "commitment to liberal democracy" for an Islamist-electoral party, or "market vs. statist" for an ideology-first movement, takes contextual judgment the rubric didn't yet pin down. The per-axis magnitudes (κ, r, mean |Δ|) are live on findings §7. Those two rubrics got the anchor-tightening pass (§3); the conflict axes reproduce as-is — a reliability floor, not validity.

Honest caveat. The pilot coders are LLM second-passes that share the project's rubric framework, so their agreement is a floor, not proof. The independent Gemini, Grok, Kimi, and DeepSeek passes — and the planned human anchors — are what make the panel a genuine cross-method check; external domain-expert validation is a v0.3 milestone.

12Status & legal taxonomy

A yes/no "in government" flag can't hold Hezbollah's coalition partnership-with-paramilitary- autonomy, the Haredi kingmakers in Israel, or the extra-parliamentary tail a binary forces into the wrong bucket. Parties carry a richer status across two independent dimensions:

  • Government role — eight tiers: lead party, major or minor coalition partner, confidence-and-supply support, major or minor opposition, extra-parliamentary, banned. Currently 10 lead parties, 28 in some form of opposition, 5 banned.
  • Legal status — legal, restricted, outlawed, dissolved, or merged-away. 5 outlawed (the Northern Islamic Movement, FIS in Algeria, the Brotherhood's political wing), 8 dissolved (Al-Wefaq in Bahrain, the Syrian Ba'ath), 0 restricted — each with a short note explaining the designation.
  • Opposition and independence flags sit alongside government role by design — a party can be neither in government nor in formal opposition (extra-parliamentary, dissolved, banned).

13Beyond positions — the surrounding corpus

The dataset is more than coordinates. The same sourcing discipline backs several connected layers:

And the standalone Knesset election machine reads the live Tayyar positions to turn vote shares into seats, governments, and a Monte-Carlo forecast — the clearest demonstration that the dataset is a substrate, not just a chart. The atlas, map, and voices gallery are alternate cuts of the same underlying data.

14Data model & reproducibility

The whole thing is a relational dataset, downloadable in full. The core tables:

  • tayyar_countries, tayyar_parties, tayyar_politicians, tayyar_party_alliances — the entities and their relationships.
  • tayyar_axes — the rubric itself: definitions, poles, scoring conventions, anchors.
  • tayyar_aggregate_positions — the published synthesis (score, lens, agreement, status, n_raters).
  • tayyar_ratings — the panel substrate: one row per (entity, axis, rater, round).
  • tayyar_analyses — document-grounded scores, per axis, with the cited spans.
  • tayyar_documents, tayyar_quotes, tayyar_events, tayyar_bills, tayyar_comparison_essays — the sourcing and surrounding corpus.

Every score, citation, and status assignment is browsable through the live tools and catalogued at the data page. A bulk export is not yet public. The instrument code (to be released under the MIT licence) and a permanent dataset release with a DOI are in preparation and will ship with the archival deposit; until then, cite the frozen tayyar-dataset-v0.2 snapshot rather than the live site.

15Limitations

The fuller account is on limitations. The ones that matter most:

  • The hand-coded priors inherit the author's reading biases. The verified-citation layer documents the facts about each entity; it does not validate the position scores. That's what the panel and grounding pipeline are for.
  • English-language sources still predominate in the citation set even where Arabic or Hebrew primaries would be more authoritative. The schema takes multiple citations per field; the document pipeline is diversifying them.
  • The choice of which axes to score is itself a political act. The 16 were chosen to surface cleavages the 2-axis compass collapses; they are not exhaustive — normalization, the peace process, climate, and others remain unaxed.
  • The compass-vs-issue split is editorial: the two compass axes are not more important, they're just the default 2-D projection.
  • Politicians inherit their party's position until individually scored; where a figure diverges from their party it's noted in prose, but quantifying it needs per-individual document scoring (on the roadmap).
  • Reasonable researchers will disagree with specific placements. Every placement is reproducible and every correction is in the migration history — disagreement is invited, not hidden.

16Roadmap — what's next

The order matters: the frame is locked first, then scaled. Document-grounded scoring is gated by the corpus, and re-scoring after an axis change is wasteful — so the axes, applicability, source kinds, and entity list are fixed before the big push, and the panel is then built once on the final frame. That way inter-rater reliability is validated a single time on the final set, not re-validated after every change.

  1. Lock the frame (under way) — axis applicability + the new context axes (§3), the social source kind (§7), the historical-party policy (§2), and a tight licence (§17).
  2. Complete the entity lists — current and historical/dissolved parties; the latter are scored in their era and marked inactive, so Ben-Gurion-era Mapai becomes comparable to today's Likud without polluting the current map. A periodic de-duplication audit flags name-variant and joint-list collisions for review.
  3. Seed to 1000+ sources — ≥3 documents per party per axis, ≥1 per politician per axis (primary, secondary, and analytic "perceived" sources), plus bios, speeches, bills, and social posts.
  4. Score fresh on the locked frame — the document-grounded panel (Claude grounded + the rounds r2–r5) is built once, in one pass, with reliability validated a single time.
  5. Make disagreement legible — contested cells render a range, not a point, and the per-rater panel page already lists every score by every rater for every multi-rater cell (live now).
  6. Findings, reproducibility, and publication — open the repository, mint a versioned dataset release (§17), and publish Paper 1 (preprint).
  7. The election machine — deepen the Israel simulator (automatic scenarios, party formations and mergers, richer forecasting) toward the coming election.

17Licence, citation & versioning

Licence & reuse

The dataset — positions, scores, ratings, axes, entity records, events, bills, quotes, briefs — is released for non-commercial research and journalism under CC BY-NC-SA 4.0: share and adapt it with attribution to "Tayyar — tarekgara.com/tayyar", a link to the exact version, and share-alike. Source texts (document excerpts, quotes, social posts) are reproduced under fair use / the right of quotation — rights remain with their original authors, and social captures carry a takedown path via contact. The code will be released under the MIT licence with the dataset deposit. Records describe public figures in their public political capacity and score the position a text takes, not the worth of a person. Full terms ship with that release.

Versioning & snapshots

The dataset is live and changes daily, so "latest" isn't citable. Each publication-grade cut is a frozen release: a tagged commit (tayyar-dataset-vX.Y) carrying full CSV + JSON exports and a manifest (row counts, schema version, generated-at date) checked into the repository, and — once the repository is public — mirrored to a permanent Zenodo DOI. Cite the tag or the DOI, never the live site. BibTeX and APA forms are on the dataset page; Paper 1 (preprint) is at /tayyar/paper.

Frozen release

Current frozen release: tayyar-dataset-v0.2 (June 2026). The citation figures below are the snapshot totals — fixed at tag time, deliberately distinct from the live counts above, which keep moving as the corpus grows.

98
parties
274
politicians
20
countries
16
active axes
286
primary-source documents
12,867
panel ratings

Documents ground scores from platforms, charters and speeches where they exist. Grounded scoring runs wherever documents exist; where a grounded reading and the prior baseline genuinely diverge, the cell is published contested (visible on the entity page and the panel page) rather than silently overwritten. The corpus keeps growing, so later cuts will carry higher coverage.

Tayyar is a living dataset in a pilot phase. It is wrong in places, thin in others, and explicit about both. What it promises is not certainty but traceability: every coordinate on every chart can be walked back to a rubric, a rater, and a source.