Piaf — The French-language dataset of Questions-Answers

Description

### [Piaf](https://piaf.etalab.studio), build an open French-language dataset for AI The use of artificial intelligence in public action is often identified as an opportunity to query documentary texts and produce automatic QR tools for users. Questioning the work code in natural language, providing a conversational agent for a given service, developing efficient search engines, improving knowledge management, all of which require a body of quality training data in order to develop Q & A algorithms. The PIAF dataset is a public and open Francophone training dataset that allows to train these algorithms. Inspired by [Squad](https://rajpurkar.github.io/SQuAD-explorer/), the well-known dataset of English QR, we had the ambition to build a similar dataset that would be open to all. The protocol we followed is very similar to that of the first version of Squad (Squad v1.1). However, some changes had to be made to adapt to the characteristics of the French Wikipedia. Another big difference is that we do not employ micro-workers via crowd-sourcing platforms. After several months of annotation, we have [a robust and free annotation platform](https://github.com/etalab/piaf), a sufficient amount of annotations and a well-founded and innovative community animation and collaborative participation approach within the French administration. ### PIAF: a shared tool of the IA Lab In March 2018, France launched its national strategy for artificial intelligence. Piloted within the Interdepartmental Digital Branch, this strategy has three components: research, the economy and public transformation. Given that the data policy is a major focus of the development of artificial intelligence, the Etalab mission is piloting the establishment of an interministerial “Lab IA”, whose mission is to accelerate the deployment of AI in administrations via 3 main activities: 1. Build a core team to internalise skills and expertise around AI 2. Supporting AI projects in administrations through calls for expressions of interest 3. Co-build shared tools that can be used as openly as possible The PIAF project is one of the shared tools of the IA Lab. ### Descriptive of the data made available The dataset follows the format of [Squad v1.1](https://rajpurkar.github.io/SQuAD-explorer/explore/1.1/dev/). PIAFv1.2 contains 9225 Q & A peers. This is a JSON file. A text file illustrating the schema is included below. This file can be used to generate and evaluate Question-Response templates. For example, following these [instructions](https://huggingface.co/transformers/examples.html#squad). ### Thanks to the 500 contributors! We deeply thank our contributors who have made this project live on a voluntary basis to this day. ### Links Information on the protocol followed, the project news, the annotation platform and the related code are here: * https://piaf.etalab.studio/ * https://piaf.etalab.studio/actualites.html * https://github.com/etalab/piaf * https://github.com/etalab-ia/piaf-code

Resources

Name Format Description Link
8 https://www.data.gouv.fr/api/1/datasets/r/14159082-d1be-417e-a67c-c3c494c7a4ad
23 https://www.data.gouv.fr/api/1/datasets/r/e32778a3-b93c-44e3-ae57-2296eb392c9e
74 https://www.data.gouv.fr/api/1/datasets/r/c0512ea0-7ed5-4aa8-9d6d-0b3f8bde021b
23 https://www.data.gouv.fr/api/1/datasets/r/063725d4-ac70-44c9-9c34-00f60d6ea9a1

Tags

  • moteur-de-recherche
  • question-answering
  • chatbot
  • nlp
  • intelligence-artificielle
  • annotation

Topics

Categories