Showing posts with label tool. Show all posts
Showing posts with label tool. Show all posts

2021-11-26

Will the color change dramatically next year ?

[中文版]

Dark olive green, not so many in 2020 as a result of COVID-19.

I am thankful that I live in Taiwan.

At the end of the year, 2021, I can’t help thinking whether the color change of many countries next year will be dramatic .

I just play a mapping via remixing "The cartogram of the World Population" and a further map layer with the Democracy Index 2020 of the Economist Intelligence .

By adding the world share of the population as part of the metadata, the map now can show the EIU score and the world share of each country.

 

The normal view:

URI: https://andrea-index.blogspot.com/2021/11/EIU-mapping.html

2017-03-08

Introduction: Metadata as Linked Data for Research Data Repositories

中文版




 “Every man has his own cosmology and who can say that his own is right.” said by Einstein. This is also true when we come to understand data semantics that one data may be differently interpreted by different data creators, curators and re-users. Then, how do we build a better research data repository?

We start with the point made by Willis, C., Greenberg, J., & White, H. (2012) that the metadata of research data increases the access to and reuse of the data. At the same time , Stanford, Harvard, and Cornell also believe the use of linked data technologies is a promising method to gather contextual information about research resources.

As a result, we look for for inspiration tools that can meet the urgent needs of innovative solutions providing feature-rich services for helping data publishing such as visualization, validation & reuse in different applications by research repositories (Assante, et.al, 2016). The CKAN (Comprehensive Knowledge Archive Network) as a major solution that makes linked metadata available, citable, and validated becomes our first choice. 


PDF 
For a general overview, our current results at data.odw.tw include:

1. An Use Case for Curation, Publication & Reuse of Metadata as Linked Data.
  • 843,309 CC licensed metadata records of 14 domains reused from the Union Catalog of Digital Archives Taiwan.
  • 44,806,400 triples (Linked Data) encoded with Dublin Core 15 Elements and Provenance Information.
  • 25,913,304 triples from 832,803 records semantically refined with spatial & temporal normalization, mapping, and linking with domain knowledges (external vocabularies, ontologies. and knowledges bases).
  • 14 domains include Archaeology,  Architecture, Archives, Artifacts, Biology, Geology, Manuscript, Multimedia, NewsMedia, PaintCal , RareBook, ResearchReuse, StoneRub.
  • 80 projects and 74 agents associated with metadata records are curated by their linked data formats and Wikidata ID: they have roles in NGO (2), Museum (5), Library (2), Government (9), Archive (1) and Academia (55).

2. A New Method to Manage Data for General-use & Discipline-specific Repositories.
  • For Open Science: using the CKAN (Comprehensive Knowledge Archive Network) as a major solution that makes linked metadata available, citable, and validated.
  • Availability: data shared with multiple formats, CSV, XML, Turtle, RDF/XML, JSON-LD, consumed both by human & machine.
  • Validation and Reproducibility: each data encoded with provenance in details while at the same time a complete mechanism for publishing article, data and code is designed and implemented.
  • A flexible and adaptable ontology for describing different data context (common knowledge or domain knowledge), event concepts (people, place, time) and objects collected by meaningful groups of different vocabularies is provided .
  • Data Visualization is enhanced and integrated through spatial and temporal mapping, filtering and linking system design.   

3. Data Semantically Enriched with Vocabularies and Knowledge Bases via Adaptable Mechanisms.
  • 18 international vocabularies used  for modeling common knowledge, and 5 domain specific vocabularies for place, time, art and humanity, or biology are applied.  3 knowledge bases like GeoNames, Wikidata, and Encyclopedia of Life are mapped and linked. 
  • The use of SPARQL language and endpoints  provide data analytic semantic queries both in local and external. In addition, data from the RDF triplestore can be easily used in 3rd-party applications.
  • Multiple DataClean Versions Mechanism:  we treat data cleaning as a kind of interpretation. Refined Versions (R Versions i.e. r1, r2, r3… ) provide different contexts to different needs of users. 
  • Multiple LinkedKnowledge Bases Mechanism:  more knowledge bases like DBpedia, WordCat, or LinkedGeoData can be linked in future via different R Versions without sacrificing  the  integrality of the original Version, encoded with DC 15 .
  • Multiple SemanticStructure Versions Mechanism:  different  interpretations results from the use of different vocabularies .  Co-exists of multiple R Versions  with different vocabularies or transforming vocabularies via SPARQL are  solutions. 
Reference:
  • Douglas, A. Vibert. "Forty minutes with Einstein." Journal of the Royal Astronomical Society of Canada 50 (1956): 99. P.100
  • Assante, M., Candela, L., Castelli, D., & Tani, A. (2016). Are scientific data repositories coping with research data publishing?. Data Science Journal, 15.. DOI: http://doi.org/10.5334/dsj-2016-006
  • Willis, C., Greenberg, J., & White, H. (2012). Analysis and synthesis of metadata goals for scientific data. Journal of the American Society for Information Science and Technology, 63(8), 1505-1520.
  • Cheng-Jen Lee, Andrea Wei-Ching Huang, Tyng-Ruey Chuang (2017) Metadata as linked data for research data repositories, International Symposium on Grids & Clouds (ISGC) 2017
  • 黃韋菁, 李承錱, 莊庭瑞 (2017) 結構資料的再次使用:語意、連結與實作, 圖書館學與資訊科學, 第 43 卷 第 1 期 2017 年 4月

Citation Information: Andrea Wei-Ching Huang (2017) Introduction: Metadata as Linked Data for Research Data Repositories. URL: https://andrea-index.blogspot.com/2017/03/metadata-as-linked-data-for-research.html

2016-10-04

LOD can be hugged by human: data.odw.tw.





LODs are nowadays all in love with machines (try 614 examples represented with 3 event types: schema:CreateAction, schema:OrganizeAction, and schema:PublicationEvent.)  Yet it can be hugged by human through CKAN at data.odw.tw
This is the Prat II of the Story of One Leaf, an implementation of the R4R ontology for reusing digital objects for 843312 reused objectsFor instance, see the leaf Pleione formosana Hayata (台灣一葉蘭) in the reusing context.


The Story of One Leaf/ D Version Examples: data:d2148340 (Pleione formosana Hayata 台灣一葉蘭)
data:d2148340 a data:Reused, r4r:RRObject, dcat:Dataset ;
r4r:hasProvenance data:p20160530-d2148340 ;
dc:publisher "中央研究院生物多樣性研究中心"^^rdf:PlainLiteral ;
dc:source "台灣本土植物資料庫
(http://taiwanflora.sinica.edu.tw/)"^^rdf:PlainLiteral ;
dc:date "採集日期:1993-04-25"^^rdf:PlainLiteral ;
dc:coverage "國家:台灣"^^rdf:PlainLiteral,
"最低海拔:1650"^^rdf:PlainLiteral,
"行政區:宜蘭縣大同鄉"^^rdf:PlainLiteral .
  • The case data:d2148340 in details is a data:Reused resource, a basic data component in the current Linked Data in the Open Data Web. It uses the R4R ontology to declare its URI as http://data.odw.tw/record/d2148340, and thus it can be shared later with its refined semantic versions using the same URI (ex. its R1 Version). It has provenance information using r4r:hasProvenance to relate the r4r:RRObject with its provenance information.
  • The data:Reused resource is a dcat:Dataset which is a collection of triples, published and curated by ODW, and available for access or download in many versions and formats.
  • Currently, the data:Reused resources are described with Dublin Core 15 Elements (DC15) in close relation to  their primary data. Here we call them DC 15 Version or D Version. Further semantic refinement is done by the refined versions, defined as data:Refined (we call them R Version, with different R1, R2, R3...) extracting information from the values of DC 15 properties. 



    See the leaf  in a semantically refined context:






The Story of One Leaf/ R Version Examples: data:d2148340 (Pleione formosana Hayata 台灣一葉蘭)
data:d2148340 a data:Refined, r4r:Data, dcat:Dataset ;
r4r:hasProvenance data:p20160706-d2148340 ;
txn:hasEOLPage eol:1134120;
dct:requires evt84:event-d2148340, evt84:phyCre-d2148340 ;
dcat:landingPage r1:r1-r2148340;
dcat:themeTaxonomy data:Biology .
evt84:phyCre-d2148340 a schema:CreateAction ;
event:factor dct:PhysicalResource ;
event:product dwc:PreservedSpecimen ;
dwc:eventDate "1993-04-25" ;
skos:inScheme dwc:HumanObservation ;
skos:scopeNote "specimen collection process" .
  • The data:d2148340 is a data:Reused, but also a data:Refined in different contexts. It is a semantically refined version (a r4r:Data in the R Version) of the data:Reused (a r4r:RRObject in D Version). In other words, data:d2148340 and r1:r1-r2148340 share the same URI according to the definition of the r4r:RRObject. which conveys the relation that the r1:r1-r2148340 (a r4r:Data) is a subclass of the data:d2148340 (a r4r:Object).
  • What distinguishes D Version and R Version most ?
    (1) the former with literal values, the latter with URI resource values.
    (2) the latter heavily relies on the event model (further described by using the former that the subject  dct:requires events to denote its people-place-time relations (extracting values of dc:coverage and dc:date from the former). Ex. the event evt84:phyCre-d2148340 triples as shown in the example.
    (3) the latter based on the curation context (data:ConceptTheme, ex. data:Biology) to enrich the resource with domain knowledge like EOL pages, dwc:PreservedSpecimenl and dwc:HumanObservation.




The following slides are a general overview for data.odw.tw.




Triples of the Girl Lost in Thought for Human:

Version D: represented with DC 15 Elements and Provenance (PROV-O)
Version R(r1): represented with refined spatial temporal information and linked to Wikidata, GeoNames, AAT …)