Skip to content
ParlaIberoPlenary debates, turn by turn

Tools

Two open tools. One lets you ask the data questions without coding. The other lets you build a corpus like this one.

Ask the data from an AI assistant

An AI assistant, such as Claude, can query the 16 corpora for you. Write the question in your own words. The assistant searches, counts, and returns numbers, passages, and charts.

It does so through an open connector, parlaibero-mcp. MCP is the standard that assistants use to work with external tools. It works with Claude, ChatGPT in its desktop app, Antigravity, OpenCode, Cursor, and others.

Install it

You need uv, a free Python installer. Then, in Claude Code, one line:

claude mcp add parlaibero -s user -- uvx parlaibero-mcp

For other programs, the repository's README (README.md) has the configuration.

View the repository

The data are downloaded from Harvard Dataverse to your computer and queried there. Your questions do not travel from your computer to us. The 16 countries take up about 26.6 GB; El Salvador, 98 MB.

What you can ask it

  • How often a word or a phrase is said, year by year, in one chamber or in several.
  • When a term first appears in each chamber, and who said it.
  • Which words set two groups apart: women and men deputies, two parties, two periods.
  • How much each group speaks compared with its weight in the chamber.
  • Passages in their context, to read them.
  • The series ready for R or Python, its chart, and the citation of each dataset.

What it warns you about

The connector knows the measured limits of each corpus and states them. It warns you when a year rests on little data, when the chair dominates a comparison, and when documents read aloud in the plenary inflate a count.

It counts and searches; it does not interpret. The reading, and citing each dataset with its DOI and its edition, are up to whoever asks.

Examples

Real questions, written the way you would write them. Below each one, what the assistant does and what it returns.

One chamber

  1. “When does the plenary in Peru talk about presidential vacancy?”

    It counts the word year by year, per million words spoken.

    A series with peaks in 2017, 2018, 2020, and 2022: the years of the vacancy motions. With its chart, ready for an article.

    ngram_viewer

  2. “In Mexico's Cámara de Diputados, who talks about parity (‘paridad’)?”

    It counts the word by sex since 2014 and divides it by everything each group says.

    Women deputies say it 74.8 times per million words; men deputies, 22.3.

    term_frequency

  3. “Which words set women and men deputies apart in Spain's Congreso de los Diputados between 2016 and 2023?”

    It compares the vocabulary of the two groups. It warns that the chair dominates the comparison, and repeats it without the chair.

    With the chair included, the vocabulary of voting comes first: votación, votos, and pausa. Without it, among the top words: mujeres, violencia, personas, and igualdad.

    distinctive_words

Several chambers

  1. “When does the word ‘feminicidio’, or ‘femicidio’, reach each chamber?”

    It searches both forms at once and dates their first appearance in each chamber, with the speaking turn that contains it.

    It reaches thirteen chambers between 2001 and 2018. If you search only for ‘feminicidio’, Guatemala almost disappears after 2009: its law says ‘femicidio’.

    term_counter · ngram_viewer

  2. “Compare how much climate change is discussed in the chambers of Spain, Portugal, and Latin America.”

    It adds up the Spanish and Portuguese forms and draws one line per chamber.

    Five chambers peak in 2019, the year of the Madrid climate summit.

    ngram_viewer

  3. “Is corruption discussed more in Argentina or in Spain?”

    Before comparing, it checks what the count rests on. It warns that in Argentina 54% of the words are in turns of more than ten thousand, almost always documents read aloud.

    Without those turns, Argentina's rate in 2010 goes from 18.6 to 31.3 per million; Spain's does not change. The honest comparison is the second one.

    coverage · ngram_viewer

First appearance of ‘feminicidio’ or ‘femicidio’ in each chamber's plenary
  1. Costa Rica2001
  2. Dominican Republic2002
  3. Panama2002
  4. Mexico2003
  5. Guatemala2004
  6. Chile2005
  7. Peru2007
  8. Colombia2008
  9. Uruguay2008
  10. Argentina2009
  11. Ecuador2011
  12. Paraguay2016
  13. El Salvador2018

Year of the first speaking turn that contains either form, in the published edition of each dataset. Measured by the connector.

Each answer carries the edition of the data it used. The assistant can also write you a methods note with everything it has run.

Build or reproduce a corpus

The method used to build the 16 corpora is published. It consists of twenty-one skills—instruction files that Claude Code follows step by step—and the code they run.

They take a chamber's Record from the scanned PDF or the HTML to ParlaIbero's 16 columns. Text recognition, cleaning, speakers, linkage with the roster, and documentation, with a log of every decision.

Who it is for

For research. Reproduce the process for a country, review a methodological decision, or add years, committees, or another chamber in the same format.

For a parliament or an organization. Turn your chamber's Records into a searchable table, one speaking turn per row, without depending on us.

View the method, step by step

It requires Claude Code, Python and, for scanned PDFs, Ollama, which recognizes the text on your computer. The skills are written in Spanish.

git clone https://github.com/rodrodr/parlaibero-tools.git
cd parlaibero-tools
bash install_skills.sh

It is the method as it was applied, in its current version; it is not a finished application. The exact copy of the code that produced each published edition is kept separately.

View the repository

The code is open, under the MIT license. If you use it in your work, cite the software as its CITATION.cff file indicates, and also cite each country's dataset with its DOI.