{"status":"OK","data":{"id":199344,"identifier":"3QTEKH","persistentUrl":"https://doi.org/10.34894/3QTEKH","protocol":"doi","authority":"10.34894","separator":"/","publisher":"DataverseNL","publicationDate":"2021-11-11","storageIdentifier":"surf://10.34894/3QTEKH","effectiveDatasetFileCountLimit":10000,"datasetFileUploadsAvailable":9993,"datasetType":"dataset","locks":[],"latestVersion":{"id":12409,"datasetId":199344,"datasetPersistentId":"doi:10.34894/3QTEKH","datasetType":"dataset","storageIdentifier":"surf://10.34894/3QTEKH","versionNumber":2,"internalVersionNumber":18,"versionMinorNumber":1,"versionState":"RELEASED","latestVersionPublishingState":"RELEASED","deaccessionLink":"","lastUpdateTime":"2021-12-23T15:41:47Z","releaseTime":"2021-12-23T15:41:47Z","createTime":"2021-12-23T06:47:25Z","publicationDate":"2021-11-11","citationDate":"2021-11-11","effectiveDatasetFileCountLimit":10000,"datasetFileUploadsAvailable":9993,"license":{"name":"CC0-1.0","uri":"http://creativecommons.org/publicdomain/zero/1.0","iconUri":"https://licensebuttons.net/p/zero/1.0/88x31.png","rightsIdentifier":"CC0-1.0","rightsIdentifierScheme":"SPDX","schemeUri":"https://spdx.org/licenses/","languageCode":"en"},"fileAccessRequest":false,"metadataBlocks":{"citation":{"displayName":"Citation Metadata","name":"citation","fields":[{"typeName":"title","multiple":false,"typeClass":"primitive","value":"Connecting Conditionals (Reuneker 2022; dissertation)"},{"typeName":"subtitle","multiple":false,"typeClass":"primitive","value":"A Corpus-Based Approach to Conditional Constructions in Dutch"},{"typeName":"alternativeURL","multiple":false,"typeClass":"primitive","value":"https://www.reuneker.nl/dissertation"},{"typeName":"author","multiple":true,"typeClass":"compound","value":[{"authorName":{"typeName":"authorName","multiple":false,"typeClass":"primitive","value":"Reuneker, Alex"},"authorAffiliation":{"typeName":"authorAffiliation","multiple":false,"typeClass":"primitive","value":"Leiden University"},"authorIdentifierScheme":{"typeName":"authorIdentifierScheme","multiple":false,"typeClass":"controlledVocabulary","value":"ORCID"},"authorIdentifier":{"typeName":"authorIdentifier","multiple":false,"typeClass":"primitive","value":"0000-0003-4206-3897"}}]},{"typeName":"datasetContact","multiple":true,"typeClass":"compound","value":[{"datasetContactName":{"typeName":"datasetContactName","multiple":false,"typeClass":"primitive","value":"Reuneker, Alex"},"datasetContactAffiliation":{"typeName":"datasetContactAffiliation","multiple":false,"typeClass":"primitive","value":"Leiden University"}}]},{"typeName":"dsDescription","multiple":true,"typeClass":"compound","value":[{"dsDescriptionValue":{"typeName":"dsDescriptionValue","multiple":false,"typeClass":"primitive","value":"Scripts (Python, R) and data (corpus data from CGN and SoNaR) belonging to the PhD dissertation 'Connecting Conditionals: A corpus-based approach to conditional constructions in Dutch' by Alex Reuneker, published in 2021. The published dissertation contains references to the data and scripts and includes a password necessary for some protected files."}}]},{"typeName":"subject","multiple":true,"typeClass":"controlledVocabulary","value":["Arts and Humanities","Computer and Information Science"]},{"typeName":"keyword","multiple":true,"typeClass":"compound","value":[{"keywordValue":{"typeName":"keywordValue","multiple":false,"typeClass":"primitive","value":"dutch, conditionals, corpus, data, scripts"}}]},{"typeName":"publication","multiple":true,"typeClass":"compound","value":[{"publicationCitation":{"typeName":"publicationCitation","multiple":false,"typeClass":"primitive","value":"Reuneker, Alex (2022). Connecting Conditionals A Corpus-Based Approach to Conditional Constructions in Dutch. (PhD dissertation. Leiden University Centre for Linguistics (LUCL), Faculty of Humanities, Leiden University). LOT Dissertation Series no. 616. Amsterdam: LOT."},"publicationIDType":{"typeName":"publicationIDType","multiple":false,"typeClass":"controlledVocabulary","value":"doi"},"publicationIDNumber":{"typeName":"publicationIDNumber","multiple":false,"typeClass":"primitive","value":"https://dx.medra.org/10.48273/LOT0610"}}]},{"typeName":"language","multiple":true,"typeClass":"controlledVocabulary","value":["Dutch","English"]},{"typeName":"grantNumber","multiple":true,"typeClass":"compound","value":[{"grantNumberAgency":{"typeName":"grantNumberAgency","multiple":false,"typeClass":"primitive","value":"NWO"},"grantNumberValue":{"typeName":"grantNumberValue","multiple":false,"typeClass":"primitive","value":"023.005.085"}}]},{"typeName":"depositor","multiple":false,"typeClass":"primitive","value":"Reuneker, Alex"},{"typeName":"dateOfDeposit","multiple":false,"typeClass":"primitive","value":"2021-11-11"}]},"dansDataVaultMetadata":{"displayName":"Data Vault Metadata","name":"dansDataVaultMetadata","fields":[{"typeName":"dansDataversePid","multiple":false,"typeClass":"primitive","value":"doi:10.34894/3QTEKH"},{"typeName":"dansDataversePidVersion","multiple":false,"typeClass":"primitive","value":"2.1"},{"typeName":"dansBagId","multiple":false,"typeClass":"primitive","value":"urn:uuid:2a502aa7-9608-4fc3-ae98-bb9914c89983"},{"typeName":"dansNbn","multiple":false,"typeClass":"primitive","value":"urn:nbn:nl:ui:13-714512bf-2ac5-4102-84a7-f9a26a49fc30"}]}},"files":[{"description":"This dataset of annotated Dutch conditionals is described in Reuneker (2022). Unfortunately, due to copyright restrictions, the actual corpus data had to be removed. The columns were kept in tact, but the actual sentences have been replaced by the placeholder '___removed due to copyright restrictions___'. Nevertheless, the dataset can still be used for inspection of feature distributions, and for clustering (see the scripts folder in this repository). Corpus locations have also been included in the dataset, which enables look-up of the texts in the original corpora, the Corpus Gesproken Nederlands (CGN) and SoNaR-500.","label":"corpus_conditionals.zip","restricted":false,"version":4,"datasetVersionId":12409,"dataFile":{"id":216498,"persistentId":"","filename":"corpus_conditionals.zip","contentType":"application/zip","friendlyType":"ZIP Archive","filesize":175848,"description":"This dataset of annotated Dutch conditionals is described in Reuneker (2022). Unfortunately, due to copyright restrictions, the actual corpus data had to be removed. The columns were kept in tact, but the actual sentences have been replaced by the placeholder '___removed due to copyright restrictions___'. Nevertheless, the dataset can still be used for inspection of feature distributions, and for clustering (see the scripts folder in this repository). Corpus locations have also been included in the dataset, which enables look-up of the texts in the original corpora, the Corpus Gesproken Nederlands (CGN) and SoNaR-500.","storageIdentifier":"surf://store:17de0f477f8-8061b2c40246","rootDataFileId":-1,"checksum":{"type":"SHA-1","value":"1f18bce92f8ec1d07f8ba9bc8e80cd4f1d7c621c"},"tabularData":false,"creationDate":"2021-12-22","publicationDate":"2021-12-22","lastUpdateTime":"2021-12-23T15:41:47Z","fileAccessRequest":false}},{"description":"These plots of feature distributions and cluster evaluations are taken from the dissertation 'Connecting Conditionals' (Reuneker, 2022).","label":"plots.zip","restricted":false,"version":2,"datasetVersionId":12409,"dataFile":{"id":216499,"persistentId":"","filename":"plots.zip","contentType":"application/zip","friendlyType":"ZIP Archive","filesize":4948134,"description":"These plots of feature distributions and cluster evaluations are taken from the dissertation 'Connecting Conditionals' (Reuneker, 2022).","storageIdentifier":"surf://store:17de0f4bbe4-e58277d4618d","rootDataFileId":-1,"checksum":{"type":"SHA-1","value":"18ebff35bda1fc8f04f516b9c81c64b0fb3d95ab"},"tabularData":false,"creationDate":"2021-12-22","publicationDate":"2021-12-22","lastUpdateTime":"2021-12-23T15:41:47Z","fileAccessRequest":false}},{"description":"These scripts can be used to automatically annotate a number of features in Dutch sentences, to convert CGN data to plain text and POS-tagged text, combine control corpora for inter-annotator agreement, convert csv to SQL for database insertion, and to convert the SoNaR corpus to plain text and POS-tagged text. Finally a set of scripts to search the plain text and POS-tagged corpus is included. The files include all necessary scripts, but not the data due to copyright. You may need additional packages to run the scripts, and as with most scripts, you might need to tweak the code to have it run on your own data.","label":"python_corpus_processing.zip","restricted":false,"version":2,"datasetVersionId":12409,"dataFile":{"id":216500,"persistentId":"","filename":"python_corpus_processing.zip","contentType":"application/zip","friendlyType":"ZIP Archive","filesize":45780,"description":"These scripts can be used to automatically annotate a number of features in Dutch sentences, to convert CGN data to plain text and POS-tagged text, combine control corpora for inter-annotator agreement, convert csv to SQL for database insertion, and to convert the SoNaR corpus to plain text and POS-tagged text. Finally a set of scripts to search the plain text and POS-tagged corpus is included. The files include all necessary scripts, but not the data due to copyright. You may need additional packages to run the scripts, and as with most scripts, you might need to tweak the code to have it run on your own data.","storageIdentifier":"surf://store:17de0f4ef34-b931a9565a2e","rootDataFileId":-1,"checksum":{"type":"SHA-1","value":"3d42ef9ffcbc504dbde55313aa8661db9d0127dc"},"tabularData":false,"creationDate":"2021-12-22","publicationDate":"2021-12-22","lastUpdateTime":"2021-12-23T15:41:47Z","fileAccessRequest":false}},{"description":"This file includes scripts that can be used to compare annotations/ratings of two annotators. See Gwet's (2014) book, of which one of the scripts is included here. Copyright remains with Kilem Gwet.\n\nThe file also includes scripts that can be used to compute inter-annotator agreement scores in a range of settings and with multiple raters. It also includes code to generate distributions of agreement scores for analysis of variance on agreement. (See chapter 4 in Reuneker, 2022 for details.)\n\nYou may need additional packages to run the scripts, and as with most scripts, you might need to tweak the code to have it run on your own data.","label":"r_classification_agreement.zip","restricted":false,"version":3,"datasetVersionId":12409,"dataFile":{"id":216501,"persistentId":"","filename":"r_classification_agreement.zip","contentType":"application/zip","friendlyType":"ZIP Archive","filesize":34006,"description":"This file includes scripts that can be used to compare annotations/ratings of two annotators. See Gwet's (2014) book, of which one of the scripts is included here. Copyright remains with Kilem Gwet.\n\nThe file also includes scripts that can be used to compute inter-annotator agreement scores in a range of settings and with multiple raters. It also includes code to generate distributions of agreement scores for analysis of variance on agreement. (See chapter 4 in Reuneker, 2022 for details.)\n\nYou may need additional packages to run the scripts, and as with most scripts, you might need to tweak the code to have it run on your own data.","storageIdentifier":"surf://store:17de0f521f5-b7a7fc07a8bf","rootDataFileId":-1,"checksum":{"type":"SHA-1","value":"33ddd293e7424af770c353121be5c73d9d8f7f20"},"tabularData":false,"creationDate":"2021-12-22","publicationDate":"2021-12-22","lastUpdateTime":"2021-12-23T15:41:47Z","fileAccessRequest":false}},{"description":"This readme file contains a description of the structure of this repository, as well as other necessary metadata.","label":"_readme.txt","restricted":false,"version":2,"datasetVersionId":12409,"dataFile":{"id":216497,"persistentId":"","filename":"_readme.txt","contentType":"text/plain","friendlyType":"Plain Text","filesize":2400,"description":"This readme file contains a description of the structure of this repository, as well as other necessary metadata.","storageIdentifier":"surf://store:17de0f424bc-a3e529ec5389","rootDataFileId":-1,"checksum":{"type":"SHA-1","value":"4e421adf18466ef0f7c48952a9a6420a1d650a74"},"tabularData":false,"creationDate":"2021-12-22","publicationDate":"2021-12-22","lastUpdateTime":"2021-12-23T15:41:47Z","fileAccessRequest":false}},{"description":"These scripts can be used to inspect the feature distributions of the corpus of conditionals reported on in Reuneker (2022). The folder includes all necessary scripts, but only one sample line of the actual corpus data due to copyright restrictions. You may need additional packages to run the scripts, and as with most scripts, you might need to tweak the code to have it run on your own data.\n\nThese scripts can be used to cluster conditionals, as described in Reuneker (2022). The folder includes all necessary scripts, but only samples of data due to large file sizes, as especially distance matrices can grow very large. For the sample data, see the data folder in the repository (so not the data folder in the scripts folder).","label":"r_features_clustering.zip","restricted":false,"version":2,"datasetVersionId":12409,"dataFile":{"id":216502,"persistentId":"","filename":"r_features_clustering.zip","contentType":"application/zip","friendlyType":"ZIP Archive","filesize":86054,"description":"These scripts can be used to inspect the feature distributions of the corpus of conditionals reported on in Reuneker (2022). The folder includes all necessary scripts, but only one sample line of the actual corpus data due to copyright restrictions. You may need additional packages to run the scripts, and as with most scripts, you might need to tweak the code to have it run on your own data.\n\nThese scripts can be used to cluster conditionals, as described in Reuneker (2022). The folder includes all necessary scripts, but only samples of data due to large file sizes, as especially distance matrices can grow very large. For the sample data, see the data folder in the repository (so not the data folder in the scripts folder).","storageIdentifier":"surf://store:17de0f55f6d-e1ce0354fecc","rootDataFileId":-1,"checksum":{"type":"SHA-1","value":"280e47097d7b634db029c15f160ced9aa41fba0f"},"tabularData":false,"creationDate":"2021-12-22","publicationDate":"2021-12-22","lastUpdateTime":"2021-12-23T15:41:47Z","fileAccessRequest":false}},{"description":"These files contain the models reported on in Reuneker, 2022. Included are the full model on a sample of the dataset, the informed model, a random model and a model with feature selection. See Reuneker, (2022, Chapter 6) for details.","label":"r_models.zip","restricted":false,"version":3,"datasetVersionId":12409,"dataFile":{"id":216503,"persistentId":"","filename":"r_models.zip","contentType":"application/zip","friendlyType":"ZIP Archive","filesize":396000881,"description":"These files contain the models reported on in Reuneker, 2022. Included are the full model on a sample of the dataset, the informed model, a random model and a model with feature selection. See Reuneker, (2022, Chapter 6) for details.","storageIdentifier":"surf://store:17de0f93685-83859453f15c","rootDataFileId":-1,"checksum":{"type":"SHA-1","value":"ad2a8fa1eb2d844ecebcea04d3a935b991784719"},"tabularData":false,"creationDate":"2021-12-22","publicationDate":"2021-12-22","lastUpdateTime":"2021-12-23T15:41:47Z","fileAccessRequest":false}}]}}}