The archive as data.
Narrative is a public good, and so is the record of what people said. Everything the site knows about its own videos is here to take: the record of every published piece, the contributors, and every transcript with the corrections applied, so the text matches the pages.
What the download holds. 250 transcripts and 429,440 words are published: they have a page here, and they are what the download holds. More of the channel is transcribed than is published here.
Download the archive (zip, 924 KB). Inside: videos.json, people.json, transcripts/<id>.txt, and a README with the schema, the corrections policy and the licence.
The two record files are also served on their own: videos.json and people.json. Every transcript paragraph with its timestamp is served as search-index.json, if you would rather have that shape; the search page now runs on a compact index built from the same pages, and falls back to that file in browsers that cannot run it.
What this is not. A court record. Auto-captions drop words and the timestamps are paragraph-level. For anything that matters, find the moment and watch it; every transcript paragraph on the site links to its second on YouTube.
The corrections policy, in one line: names and terms are corrected, speech is not. The ums stay, the false starts stay, and what a person actually said stays even when it is wrong. Where YouTube masked a word in its own captions it shows as [bleeped]; that is YouTube’s, not ours. The full policy is in the README.
Licence. The text data is released under Creative Commons Attribution 4.0. Use it, quote it, build on it, and credit We Them Media with a link. The videos themselves remain on YouTube under YouTube’s terms. The people in the interviews said these things on camera in public; treat them as people.
Questions about the data, or a correction: wethem.xyz@gmail.com.
