Synthetic data - top diabetes
Description
### Description of the database:
- **Objectives and initial purposes of the database:**
This synthetic dataset was created as part of the translation and implementation of the algorithm used by the CNAM to build the top diabetes ([link to the description sheet of the algorithm](https://www.health-data-hub.fr/library-open-algorithms-health/algorithm-to-build-the-top-diabete-of-mapping)).
The Python and SAS versions adapted by the HDH cover synthetic data for the years 2018-2019 but can be extended to other years. The CNAM source program was developed in SAS and runs on data from 2015 to 2019.
The objective of the algorithm mentioned above is to target people in care for diabetes in the main base of the NSDS in order to create the ‘Top Diabetes’ of the pathology mapping created and maintained by the CNAM (version G8).
- **Context of creation:**
The implementation of the top diabetes algorithm required the mobilization of synthetic (fictitious) tables and variables.
-merge annual tables into a single table for ER_PRS_F, ER_ETE_F, ER_PHA_F,
Data/SNDS community.
- **Results associated with the creation of the database:**
The algorithm used by the CNAM to construct the top diabetes: (source version (CNAM), Python version and SAS version (HDH)) (https://www.health-data-hub.fr/library-open-algorithms-health/algorithm-to-build-the-top-diabete-of-mapping).
Recruit people, in a wide variety of fields, to work in Quebec, which is looking to recruit in the region.
- **Collection methodology and inclusion criteria:**
### Data presentation:
The programmes operate on the synthetic data of the HDH with some adaptations:
This dataset was generated using the [scheme](https://gitlab.com/healthdatahub/se-former-au-snds/synthetic-generator/-/tree/master/src/resources/schemas?ref_type=heads) of the 2019 NSDS main database tables.
- **Target audience:**
-the conversion of the date format to yymmdd10.
Patient identification is based on the targeting of specific medicines and/or ALD and/or hospitalisation in MCO.
-the renaming of NUM_ENQ to BEN_NIR_PSA, The mapping algorithms aim to maximize specificity (not sensitivity), i.e. to ensure the absence of non-diabetics among the targeted patients.
- **Choice of variables:**
The implementation of the algorithm requires the mobilisation of the following tables and variables (the required history is indicated in the corresponding box):
Patients with less than 3 dispensings of specific drugs, who do not have ALD and who have not been hospitalized within 5 years for diabetes are not retained.
The programs adapted in SAS and Python run on synthetic data from the years 2018 and 2019. The CNAM source code (in SAS) was designed to work on data from the years 2015 to 2019.
### Limits of this dataset:
 the lack of medical consistency, the lack of updating of annual changes, an evolutionary table scheme that can be incomplete and imperfect.
This programme does not include an analysis of the estimated items of expenditure reimbursed by Health Insurance.
The algorithm identifies prevalent patients with diabetes in a given year (2019). It does not determine the exact date of onset of diabetes in the base.
The use of synthetic data, although useful for manipulating NSDS data, has limitations:
More information on the use of the database in the context of the top diabetes programmes (CNAM) on the GitLab repository of the programmes ([link of the GitLab repository](https://gitlab.com/healthdatahub/boas/cnam/top-diabete)).
### Support:
Contact point: [dir.donnees-SNDS@health-data-hub.fr](https://health-data-hub.fr)
### Contribution:
On Gitlab (make a ticket or merge-request)
Resources
| Name |
Format |
Description |
Link |
|
57 |
|
https://www.data.gouv.fr/api/1/datasets/r/96ca92c8-caac-424b-8f2d-1c35a8888d69 |
Tags
- cnam
- france
- health-data-hub
- cartographie
- hdh
- programme
- python
- ciblage
- donnees-synthetiques
- sas
- algorithme
- pathologies
- diabete
- version-g8