Methodology
What is in each row, what was measured, and what is not claimed.
This page does not follow the process stage by stage. It answers what a reviewer would ask, with each limit in the same sentence as its number.
Each section ends with something you can check for yourself.
ParlaIbero is a data infrastructure for the social sciences and humanities. PELA-USAL surveys parliamentarians in Latin American countries with a standard questionnaire, adapted to each country. Latinobarómetro applies the same public opinion study across the countries of the region. ParlaMint publishes debates of European parliaments with a common encoding. The Manifesto Project codes parties' election programs with a common scheme. Here the material is what was said in the plenary: the same row and the same columns in all sixteen chambers.
As of September 21, 2026, we have not located another corpus covering several Latin American countries with the full text and each speaker identified as a deputy. There are ParlaMint, which covers Spain and Portugal; ParlSpeech, which covers Spain; and ParlEE. In the region, rosters without text, presidential speeches, and national corpora.
What a row is
The unit is the speaking turn, as the Record itself marks it. Each turn is a row.
Each row carries the speaker's designation exactly as printed, the full text, the session, and its position within it. If the speaker could be identified, it also carries their identifier in the roster, with name, sex, party, and district. There are 16 columns, the same in all sixteen chambers.
id_session- SV0010058
id_int- SV001005800070
legislature- 2018-2021
legislative_session- AÑO LEGISLATIVO 2020
session_number- 60
date- 2019-07-24
session_type- ordinaria
intervention_order- 70
speaker_rawid_depspeaker_namesex- M
party- PCN
district- Sonsonate
dm_speech- 1
textGracias señor presidente, compañeros y compañeras diputados, también para unirme a la celebración digamos de aquella ciudad que de solo mencionarlo nos recordamos de El Principito, cosa que no muchos salvadoreños saben, que Armenia, una mujer de Armenia de apellido Suncín, fue la inspiración de la Rosa de El Principito. Así que me uno a la celebración como dije de Armenia, gracias […]
Who speaks is in the data. It is not shown here.
The text continues: only the beginning is shown here.
The text is neither summarized nor lemmatized. The order is preserved. And the order is information.
The Record is not the session: it is what each chamber published of it. The median ranges from 11 words per turn in Uruguay to 92 in El Salvador. That distance measures how each chamber transcribes, not how much debate there is.
Try it. Download El Salvador, the smallest file (98 MB; it covers 2018 to 2025), and open it in the explorer. Download the data · Open the explorer
The explorer's interface is in Spanish only today. Search runs on each chamber's original text, in Spanish or Portuguese.
From sixteen typographic traditions to one table
Each chamber prints in its own way who takes the floor: “El señor APELLIDO:” in one, “O SR. NOME (partido - estado) –” in another. Five steps bring them into a single table.
- From the source to the faithful text. Digital PDF, scanned page, HTML, XML, or Word: first the text is extracted without touching anything. Then page headers, page numbers, and words split across lines are cleaned.
- From the text to the turns. Each speaker marker is located with the formulas of that chamber and that period. An unrecognized marker buries a turn inside the previous one, and no count sees it. That is why a single one stops the tagging of that session.
- From the speakers to the deputies. Each printed designation is checked against the roster: 37,523 people in 70,570 term segments, across the sixteen chambers.
- From the roster to the common schema. The row receives party, district, and sex according to the term segment that covers its date.
- From the schema to the deposit. Each country is published with 17 files—data, roster, and documentation—generated by a single program, never by hand.
The rows are produced by deterministic code. The language model wrote that code and recognized the text of the oldest scanned pages.
Try it. On your country page, see which source and which format its text came from. View the countries
Speech and non-speech
Nothing is deleted from the Record.
The cover page, the table of contents, and the roll call are kept in place, as the first row of the session: 51,741 rows across the sixteen chambers.
They carry the mark dm_speech = 0. So do the roll-call vote tallies, the narration, and the documents read aloud, where the Record itself proves that they are not speech: 222,975 rows in total.
The mark is asymmetric, on purpose. A 0 is set only where it is proven. A 1 does not mean “verified speech”: it means “not proven to be non-speech.”
The consequence is declared. In Argentina, Uruguay, Mexico, Costa Rica, Ecuador, the Dominican Republic, and Guatemala, part of what is read aloud remains inside the turn of whoever reads it. There, the word count of the chair is inflated.
Filtering is up to you. Deciding for you is not up to us.
Try it. In the explorer, the first entry of each session is “Encabezado y sumario de la sesión” (Session heading and table of contents). The roll call is there. Open the explorer
The explorer's interface is in Spanish only today. Search runs on each chamber's original text, in Spanish or Portuguese.
Who speaks
Of the speech turns, 87.05% carry an identified deputy: 8,590,210 of 9,868,087. This is gross linkage.
Effective linkage is 98.12%. It excludes two kinds of turns from the denominator, and only two. The turns of those who cannot hold a seat: ministers, administrative clerks, reading clerks, guests. And those the Record makes unattributable, such as “Varios señores diputados”.
It does not exclude an unnamed chair. Whoever chairs is a deputy: if their identity was not recovered, that is our missed link.
That leaves 1,277,877 speech turns without a deputy. Of these, 962,878, or 75.35%, belong to those who cannot hold a seat. Another 150,714, or 11.79%, are collective or anonymous voices. The remaining 164,285, or 12.86%, are our missed links.
The distance between the two rates is a property of the Record, not of the processing. In Panama, 30.68% of the speech turns belong to those who cannot hold a seat, above all the Secretaría, which reads aloud; in Uruguay, 0.81%. That is why comparison between countries starts from the table by chamber, not from the pooled number.
| Country | Chamber | Speech turns | Gross linkage | Effective linkage | Those who cannot hold a seat |
|---|---|---|---|---|---|
| Portugal | Assembleia da República | 1,554,381 | 85.54% | 98.32% | 4.14% |
| Spain | Congreso de los Diputados | 558,940 | 96.85% | 99.41% | 2.57% |
| Ecuador | Cámara Nacional de Representantes / Congreso Nacional / Asambleas Constituyentes / Asamblea Nacional | 1,280,822 | 68.76% | 94.78% | 27.46% |
| Argentina | Cámara de Diputados de la Nación | 256,636 | 90.78% | 97.16% | 6.56% |
| Uruguay | Cámara de Representantes | 421,284 | 96.41% | 97.68% | 0.81% |
| Mexico | Cámara de Diputados | 630,937 | 96.18% | 98.21% | 2.07% |
| Paraguay | Cámara de Diputados | 360,201 | 76.59% | 93.91% | 18.44% |
| Chile | Cámara de Diputadas y Diputados | 556,679 | 95.07% | 99.57% | 3.37% |
| Costa Rica | Asamblea Legislativa | 375,730 | 95.66% | 99.79% | 4.14% |
| Peru | Congreso de la República | 849,499 | 88.74% | 99.02% | 9.86% |
| Panama | Asamblea Nacional (Legislativa hasta 2004) | 508,324 | 67.93% | 98.00% | 30.68% |
| Colombia | Cámara de Representantes | 463,201 | 69.36% | 96.79% | 28.35% |
| Guatemala | Congreso de la República | 251,720 | 96.38% | 99.80% | 3.43% |
| Dominican Republic | Cámara de Diputados | 101,381 | 98.50% | 98.88% | 0.38% |
| Brazil | Câmara dos Deputados | 1,649,698 | 98.04% | 99.10% | 1.07% |
| El Salvador | Asamblea Legislativa | 48,654 | 98.38% | 99.94% | 1.56% |
Try it. Find your chamber: the two rates, and what each one excludes. Download the figure data
Sex is a derived variable
No Record states the sex of whoever speaks. The sex column is derived in the roster: from an official registry where one exists; otherwise, from the given name or the form of address. Each value carries its provenance in sex_source.
Its accuracy is 98.98%, measured against the exhaustive human review of eleven complete rosters, 38,372 rows. For men, 99.54%; for women, 97.43%. Error is 5.6 times more likely for women. The number characterizes the procedure, not each corpus.
The other five—Argentina, Brazil, Chile, Ecuador, and Peru—are not audited, and they are drawn differently in the figures. They are not alike. Brazil, Chile, and Peru take sex from an official registry; in Ecuador, every value is inferred from the name or the form of address. Anyone who needs precision restricts sex_source to manual, official_registry, and given_name; that excludes all of Ecuador.
The gender audit
Extraction defects are not neutral. When a pattern fails, the person who disappears is, more often, a woman.
In Costa Rica, the pattern recognized “PRESIDENTE” and not “PRESIDENTA”. The checker had inherited the same lexicon and reported zero residual markers. The totals matched. The turns of the women who chaired were left inside the previous turn.
In Brazil it could be measured when the pattern was corrected, before the corpus was processed again. Women accounted for 8.7% of the turns and 47.8% of the recovered turns.
The pattern is written first in the masculine. The feminine forms are more varied, and the line break splits them more often. The generic “Presidente” hides whoever chairs. Hence the rule: “PRESIDENTA” cannot be a man; “PRESIDENTE” says nothing.
Not everything went in the same direction, and we say so: in Argentina, one correction moved rows from women to men. The claim that stands does not depend on direction: each correction brought the corpus closer to the Record.
Try it. Download the data for the opening figure: words, turns, and women speakers, by chamber and decade. Download the figure data
Optical character recognition (OCR): what failed
All of Ecuador is scanned paper. So are the earliest years of Uruguay, Panama, and Paraguay, and seven Dominican sessions.
A vision model read some of those pages and sometimes wrote its own text inside the Record: its instructions, comments, and, worst of all, Spanish lines translated into English.
We removed 1,949 fragments in Uruguay, Ecuador, Panama, and the Dominican Republic. The class is bounded, not closed: the model improvises a different wording each time, and what it translated cannot be repaired by deleting.
Where the scan cropped the margin, nothing was reconstructed. What is not in the pixels is not invented.
Try it. Read the declared limitations of Ecuador, the only corpus scanned from beginning to end. View the page for Ecuador
Validation and human review
There are two different things here, and each carries its full label.
Leakage of the PASS/FLAG filter: 0.5% [0.09%–2.78%].
The sample: 300 random rows in each of fifteen countries. A deterministic filter separates the trivially correct rows—PASS—from the rest—FLAG—which goes to human review. Leakage is what the filter passes and should not: 1 of 200 rows passed and reread by hand.
It was run on July 31, 2026, before the corpus was processed again to produce the published edition, and without Ecuador. It is not the error rate of the file you download. It does not support the phrase “all sixteen validated.”
The second thing: we read the sixteen files row by row, before and after processing them again. That reading found what no check had seen: page headers and page numbers embedded in the middle of sentences, speaker markers split one word per line.
“Reviewed” does not mean “free of known defects.” The known ones are measured and declared, country by country.
Try it. Open the known limitations of your chamber. View the countries
What the corpus does not claim, and how to compare
This edition delivers the material evidence: what was said, structured, attributed, and measured. It does not include topic, tone, ideological position, or vote.
It does not claim that a dm_speech = 1 is verified speech, nor that the chair is identified in every row, nor that the scanned decades are free of errors.
The legislature column is not comparable across countries: it is what each Record prints, whether a constitutional term, a year, or a half-year. The comparable keys are id_session and the date.
How to compare: rates within a country, against itself, over time. Never volumes across countries. And the denominator, in plain view.
Coverage is a census of what each chamber published, with the gaps named: 527 of 800 chamber-year cells between 1976 and 2025. Of 21 countries, five are missing, and not at random.
Venezuela, Cuba, and Nicaragua do not publish the records of their debates. Bolivia publishes them without digitizing them. Honduras publishes summaries of what was discussed, not the interventions.
Try it. Read the data dictionary: it says, column by column, what is comparable and what is not. Read the dictionary
Editions, identifiers, and reproducibility
Anyone who builds on someone else's data needs to know what moves and what does not. Each country has its DOI in Harvard Dataverse. A published edition does not change: anything new comes out under another number.
Of the sixteen datasets, fifteen are at edition v2.0; Peru (v1.0).
The identifiers id_session and id_int are stable within an edition, not across editions. They are derived from position, and adding a session shifts those that come after. Cite the edition.
To cite a passage, give the date and the session number, and check it against the official Record. Derived edition for research: in case of any discrepancy, your chamber's Record prevails.
Reproducibility has a boundary, and it is measured. From the tagged text onward—linkage, merge, schema, and documentation—everything is deterministic. This was checked by running the sixteen processing chains twice and comparing the files.
What comes before—OCR, extraction, cleaning, tagging—falls outside that guarantee. The recognized pages and the human linkage decisions are stored frozen, as input. Anyone who cites the reproducibility of this corpus must say from where.
The full method, with its code, is published so that it can be reproduced or extended. View the tools
Try it. Choose your country and copy its citation, with its edition and its DOI. View the countries
The full documentation
This page summarizes. In full:
- With each country, in English, Spanish, and Portuguese: the README, the data dictionary, the known limitations, and the processing report. View the countries
Try it. Start with the data dictionary, right here, and continue with the El Salvador file. Read the dictionary · Download the data