Creedibility Corpus with several datasets (Twitter, Web database) in French and English

Description

Description of the corpora The set of these datasets are made to analyse ifnormation credibility in general (rumor and disinformation for English and French documents), and occuring on the social web. Target databases about rumor, hoax and disinformation helped to collection obviously misinformation. Some topic (with keywords) helps us to made corpora from the micrroblogging platform Twitter, great provider of rumors and disinformation. 1 corpus describes Texts from the web database about rumors and disinformation. 4 corpora from Social Media Twitter about specific rumors (2 in English, 2 in French). 4 corpora from Social Media Twitter randomly built (2 in English, 2 in French). 4 corpora from Social Media Twitter about specific rumors (2 in English, 2 in French). Size of different corpora: Social Web Rumorous corpus: 1,612 French Hollande Rumorous corpus (Twitter): 371 French Lemon Rumorous corpus (Twitter): 270 English Pin Rumorous corpus (Twitter): 679 English Swine Rumorous corpus (Twitter): 1024 French 1st Random corpus (Twitter): 1000 French 2nd Random corpus (Twitter): 1000 English 3rd Random corpus (Twitter): 1000 English 4th Random corpus (Twitter): 1000 French Rihanna Event corpus (Twitter): 543 English Rihanna Event corpus (Twitter): 1000 French Euro2016 Event corpus (Twitter): 1000 English Euro2016 Event corpus (Twitter): 1000 A matrix links tweets with most 50 frequent words Text data: _id: message id body text: string text data Matrix data: 52 columns (first column is id, second column is rumor indicator 1 or -1, other columns are words value is 1 contain or 0 does not contain) 11,102 lines (each line is a message) Hidalgo corpus: lines range 1:75 Lemon corpus: lines range 76:467 Pin rumor: lines range 468:656 Swine: lines range 657:1311 random Messages: lines range 1312:11103 Sample contains: French Pin Rumorous corpus (Twitter): 679 Matrix data: 52 columns (first column is id, second column is rumor indicator 1 or -1, other columns are words value is 1 contain or 0 does not contain) 189 lines (each line is a message)

Resources

Name Format Description Link
39 https://www.data.gouv.fr/fr/datasets/r/e64044af-a3bf-4be5-8295-ac6811fd72a5
39 https://www.data.gouv.fr/fr/datasets/r/685b5bee-1215-4c38-b599-438657116522
39 https://www.data.gouv.fr/fr/datasets/r/b2a188f6-e618-4c42-a5c5-59002f986d3f
39 https://www.data.gouv.fr/fr/datasets/r/554da648-f51f-4418-91f1-1b7d2a7a7113
39 https://www.data.gouv.fr/fr/datasets/r/7f14efd9-0016-4b45-9ff5-519f70e4dcf9
39 https://www.data.gouv.fr/fr/datasets/r/0cd1cc7e-5ea3-42f6-a7b6-a5df2790199e

Tags

Topics

Categories