academic · dataset overview
Tayyar dataset
A position dataset for MENA political actors. 98 parties and 274 politicians across 20 countries, scored on 16 axes. Every fact carries an external source citation. The dataset will be released as a versioned, openly-licensed snapshot with a DOI when the data paper lands; for now it's browsable through the live tools. This page is the canonical citable surface — the cards below summarize what's in the dataset, the methodology buttons jump to the deep dives.
What's included
- Field-level source citations. Each fact about each entity links back to where it came from — founding year cites a different source than current leader, which cites a different source than legal status. The coverage page tracks the rollup.
- 16 calibrated axes. Economic, social, state-religion, democracy, west-alignment, regional-stance, Palestinian question, civil liberties, regime stance, pan-Arab, federalism, modernization, gender, iran-posture, press-freedom, sectarianism. Each one comes with a scoring rubric and concrete MENA examples anchored along the scale.
- Richer status than yes/no. Parties carry a government role (lead, coalition major / minor, confidence-and-supply, opposition major / minor, extra-parliamentary, banned) and a legal status (legal, restricted, outlawed, dissolved, merged away). Opposition and independent flags sit alongside.
- Declared vs. behavioral on key cases. Parties like Hezbollah and Hamas read as more committed to democracy in what they declare than in what they do; 20 parties carry a rhetoric-vs-record gap that is itself the finding. The home compass has a lens toggle to switch between views.
- 426 primary-source documents and 101 verified quotes. The document corpus carries verbatim manifestos, charters, parliamentary speeches, and UN addresses with country / party / politician attribution. The quote corpus drives Who-said-it and the "On the record" sections on every party / politician page.
- A semi-live event feed. Pulse tracks recent political shifts with a confidence rating — confirmed, reported, rumored, speculative — so in-flux developments (party formations, merger talks) can sit alongside confirmed events without being conflated. Subscribable as RSS.
- Versioned release in preparation. The full dataset will be published as an openly-licensed, DOI'd snapshot alongside the data paper; until then it's browsable live and available from the author on request.
Cite as
Cite the paper. The dataset is in active development; a tagged snapshot pins a citation that won't shift as it evolves, and a public repository and permanent DOI are in preparation.
How to cite 1 reference
Paper 1 Gara, T. (2026). The Model as One Rater Among Several: Measuring Political Positions in Data-Sparse Regions with a Language-Model Panel. Preprint, arXiv:2606.23042.
Show BibTeX
@misc{gara_tayyar_2026,
author = {Gara, Tarek},
title = {The Model as One Rater Among Several: Measuring Political Positions in Data-Sparse Regions with a Language-Model Panel},
year = {2026},
eprint = {2606.23042},
archivePrefix = {arXiv},
primaryClass = {cs.CY},
doi = {10.48550/arXiv.2606.23042},
url = {https://arxiv.org/abs/2606.23042}
} Gara, T. (2026). The Model as One Rater Among Several: Measuring Political Positions in Data-Sparse Regions with a Language-Model Panel. arXiv:2606.23042. https://arxiv.org/abs/2606.23042
Where to read more
- Methodology — how the dataset got built and where it falls short
- Findings — the structural patterns the data shows
- Coverage — verification status, country breakdowns, special-status leaderboards
- Axes catalog — all 16 axes with correlations and per-axis stats
Data access
Bulk downloads aren't public yet. The dataset is still in active development (new scoring
rounds, a verification pass), so rather than hand out a snapshot that shifts under you, the
full data will be released as a single versioned, openly-licensed
(CC BY-NC-SA 4.0), DOI'd snapshot when the data paper lands,
with the instrument code open-sourced under MIT. Until then, every score is browsable
through the live tools with its sources cited per entity, and
the frozen v0.2 snapshot behind the paper is available from the author on request.
What's coming next
Changes are tracked as versioned snapshots; the repository will be open-sourced under MIT on publication. The roadmap, in the order it'll land:
- Document-grounded scoring. Positions derived from reading party platforms, speeches, and voting records — with the specific passages cited. The hand-coded scores stay as the baseline; the document-grounded ones replace them as they're produced.
- Inter-rater agreement. Cohen's κ between hand-coded and document-grounded scores reported on methodology. Where they agree, the rubric's doing its job; where they don't, that's a finding worth writing up.
- Lens system at scale. Declared / behavioral / perceived rows generated for every party, not just the hand-coded marquee cases. The compass lens toggle then carries information across the whole dataset.
- Second-pass verification. Each fact and each position score reviewed against primary sources by someone other than the author.