Zobrazují se příspěvky se štítkemwikipedia. Zobrazit všechny příspěvky
Zobrazují se příspěvky se štítkemwikipedia. Zobrazit všechny příspěvky

úterý 16. května 2023

Summary evaluation for Wikipedia

Wikipedia articles, at least in English, tend to be overgrown - they contain a lot of information of mixed importance. However, we do not always have time to go thru all the content. It helps that articles are structured to have the most important things in the first sentence/paragraph. However, the importance is not really differentiated within the body. If you have to read the body, you get swamp. I use two tricks to deal with that: 1. Switch to a different language. The idea is that articles in different languages are smaller. However, they still contain the most important information. 2. Use a historical version of the article. The idea is that the most important information was entered before the less important information. People are obsessed these days with text-generative AI. Hence a proposal to use AI for shortening of English articles. Do you need a short description? Generate just a single sentence. Was it not enough? Generate the rest of the paragraph. Need even more? Write a subtopic, which interests you. How to evaluate the quality of the summaries? A. Machine translate all different language variants of the article into English and check the information overlap between the summary and the language variants. Ideally, the overlap will be large. This exploits trick #1. B. Check the overlap between the summary and historical versions of the article. Ideally, the information in the summary will be present even in the old versions of the article. This exploits trick #2. Limitations: 1. Some important information is known only from some date. For example, election results are not available before the results are announced. This can be corrected by observing how quickly given information spreads across different language versions. If the information spreads quickly, it is likely important information, even though it is young information. 2. Language variants are highly correlated because they copy from each other. However, it is reasonable to assume that, for example, English and Spanish are more correlated than, for example, Tuu and Thai, simply because fewer people speak both Tuu and Thai than English and Spanish. If the compensation of these differences is necessary, estimate a correlation matrix on the data and use it to weight the signal.

středa 18. listopadu 2020

What is wrong with Wikipedia

I like Wikipedia. But I am worried about the future of Wikipedia. Why? Because it keeps growing without limit.

When Wikipedia was based, it got one thing right: the delivery of the information has to snowball. You begin with the minimal quantum of information that is self-sustainable. And then you keep iteratively expanding that. A good article at Wikipedia starts with a self-sustainable sentence, which can exist alone and provides basic description of the keyword. Then there is the rest of the first paragraph, which slightly expands the first sentence. And as the whole, the first paragraph is self-sustainable. Then there is the rest of the paragraphs that make the header. They slightly expand the description of the first paragraph. And together with the first paragraph, they are self-sustainable. And finally, there is the body, which provides the rest of the information and which is, by the fact that it completes the whole article, also self sustainable.

A good metaphor to this concept is the understanding of a picture. With the snowball approach, you first look at the picture from a distance. And you recognize a house. Then you move closer and recognize the front door and windows. You move even closer, and recognize individual parts of the doors. And finally, when you move the closest, you see details like the cracks on the door panels.

In contrast, with pinhole approach you scan the picture pixel-by-pixel. And you have to reconstruct the whole picture your mind.

While the pinhole approach is perfectly fine for computers, humans generally prefer the snowball approach. If nothing else, it allows them to skip irrelevant information like details of the clouds, because from the previous step they already know that that patch of pixels are clouds.

For long time, Wikipedia followed the snowball approach. But now, it keeps shifting to pinhole approach as the bodies of the articles keep growing without any limit.

For me, the current transition from the header to the body is frequently too abrupt. I cope with that by switching from English version of the article to some non-English version, where the articles are smaller. Most of the time, the smaller version provides all the information that I need. But even if it does not answer everything, I at least know which information I seek. And I can then quickly jump to the relevant parts in English version of the article.

But how can the situation be systematically rectified? There are multiple options:
    1) Identify an optimal size of the articles and start truncating the overgrown articles.
    2) Allow fast time traveling to time when the article was closest to the optimal size.
    3) Expandable TOC. Now, TOC doesn't even fit whole screen. Could we by default hide the lowest level headers but on click expand them?
    4) Inform Wikipedist that the article is too long and that they should abstain from making it longer. If something, they should make it shorter.
    5) Introduce another layer of granularity.

The first approach is politically unacceptable. Wikipedists like to expand articles, not to reduce them. The second approach is meaningful on topics that do not evolve rapidly, like interpretation of ancient events, but fails miserably on contemporary topics. The third approach is a nice technical solution. The forth approach might help a bit. But the last option is the only real solution that might have a chance to succeed.


pondělí 28. ledna 2013

Proč anglická Wikipedia stagnuje

Počet přibývajících článků na anglické Wikipedii má podobu klasické nasycovací S-křivky a totéž platí o počtu oprav. A Kaggle vyhlásilo soutěž na zjištění příčiny.


Dovolil bych si ale odhadnout příčinu a její možné řešení i bez nahlédnutí do dat. Stávající obsah dat nebo její forma reprezentace se nasytili. Wikipedia je založena na textovém obsahu. Jakmile ale článek nakyne do obézní velikosti, lidi přestanou být motivováni ho rozšiřovat. Naopak by ho raději viděli kratší. Pokud jste se ale někdy pokoušeli zkrátit 100 stránkovou studii na stránku, abyste ji mohli publikovat, víte, že zkracování článku je obtížný problém. A tak lidi raději nechávají články tak, jak jsou.

V případě multimediálního obsahu ale Wikipedia nenabyla nasycení. Problémem je spíše pracnost vložení multimediálního obsahu. Jak například přidám obrázek do pravého horního rohu. Jistě je na to šablona, ale kde ji najdu? Jak ji použiji? A už to začíná být složité. Jako řešení bych viděl přidání placeholderu do článků bez jediné fotografie, který by říkal: "Buďte první, kdo přidá fotografii". Po kliknutí by se objevil dialog pro nahrání fotografie z disku. Po nahrání by se ještě objevil formulář pro vyplnění důležitých metainformací, jako zda jste majitel. A Wikipedia by měla novou fotografii.

Věřím, že tenhle přístup by měl úspěch. Když člověk vidí, že článek není kompletní, je motivován ho doplnit, když může. Nahrání fotografie je jednoduché. To zná z facebooku. A doplnění metainformací? Když už se dostal až sem, tak se nevzdá na nějakém formuláři a vyplní ho. Navíc díky tomu, že se umožní jen nahrávání z počítače, tak lidi budou motivováni nahrávat jen originální fotografie, protože nahrání fotografie z internetu by bylo obtížné. Najít fotografii na Googlu, stáhnout, nahrát a nakonec ještě vyplnit formulář. Pochopitelně by byla potřeba kontrolovat, že data opravdu nejsou z internetu. Na to ale stačí automatický dotaz na Google. Když Google najde na internetu hodně podobných fotografií a nahrávač nevyplnil podrobné informace o autorství, pravděpodobně se jedná o pro Wikipedii nepoužitelnou fotografii.

U audia by byl postup podobný. Je daný článek o hudbě? Tak šup tam s audio placeholderem, ať lidi nahrávají. U skladeb starších 150 let a vlastních interpretací by to neměl být problém. Obdobně u videa nebo ontologických tagů. Zobrazte, že tam nějaká informace chybí, a někdo ji vyplní.

Z jiného soudku: občas se mi na Wikipedii stane, že pochybuji o správnosti uvedené informace. Ale nedaří se mi nikde najít informace potvrzující nebo vyvracející moji hypotézu. A tak to nechám být. Přitom bych se ale strašně rád podělil o mých pochybnostech. Psaní komentáře je složité a ponižující. Kdo by se taky veřejně hlásil k tomu, že je debil, že nechápe tak evidentní věc? Místo toho navrhuji implementovat obdobu funkcionality na Brittanice. Člověk pochybující o zobrazené informaci by ji probarvil, objevila by se kontextová nabídka a uživatel by zaškrtl: "navrhnout k revizi". A pilní wikipedisté by potom procházeli nejčastěji označovaná místa a opravovali je. Ať už opravou chyby, změnou formulace, přidáním vysvětlení nebo reference.

Jinak řečeno. Až uživatelům dáte prostor k vylepšování Wikipedie, rádi pomůžou, jako již dříve pomohli.

Update: v prosinci 2013 jsem zaznamenal, že na http://cs.wikipedia.org/wiki už začaly používat obrázkový placeholder: