Showing posts with label context. Show all posts
Showing posts with label context. Show all posts

2017-03-08

Introduction: Metadata as Linked Data for Research Data Repositories

中文版




 “Every man has his own cosmology and who can say that his own is right.” said by Einstein. This is also true when we come to understand data semantics that one data may be differently interpreted by different data creators, curators and re-users. Then, how do we build a better research data repository?

We start with the point made by Willis, C., Greenberg, J., & White, H. (2012) that the metadata of research data increases the access to and reuse of the data. At the same time , Stanford, Harvard, and Cornell also believe the use of linked data technologies is a promising method to gather contextual information about research resources.

As a result, we look for for inspiration tools that can meet the urgent needs of innovative solutions providing feature-rich services for helping data publishing such as visualization, validation & reuse in different applications by research repositories (Assante, et.al, 2016). The CKAN (Comprehensive Knowledge Archive Network) as a major solution that makes linked metadata available, citable, and validated becomes our first choice. 


PDF 
For a general overview, our current results at data.odw.tw include:

1. An Use Case for Curation, Publication & Reuse of Metadata as Linked Data.
  • 843,309 CC licensed metadata records of 14 domains reused from the Union Catalog of Digital Archives Taiwan.
  • 44,806,400 triples (Linked Data) encoded with Dublin Core 15 Elements and Provenance Information.
  • 25,913,304 triples from 832,803 records semantically refined with spatial & temporal normalization, mapping, and linking with domain knowledges (external vocabularies, ontologies. and knowledges bases).
  • 14 domains include Archaeology,  Architecture, Archives, Artifacts, Biology, Geology, Manuscript, Multimedia, NewsMedia, PaintCal , RareBook, ResearchReuse, StoneRub.
  • 80 projects and 74 agents associated with metadata records are curated by their linked data formats and Wikidata ID: they have roles in NGO (2), Museum (5), Library (2), Government (9), Archive (1) and Academia (55).

2. A New Method to Manage Data for General-use & Discipline-specific Repositories.
  • For Open Science: using the CKAN (Comprehensive Knowledge Archive Network) as a major solution that makes linked metadata available, citable, and validated.
  • Availability: data shared with multiple formats, CSV, XML, Turtle, RDF/XML, JSON-LD, consumed both by human & machine.
  • Validation and Reproducibility: each data encoded with provenance in details while at the same time a complete mechanism for publishing article, data and code is designed and implemented.
  • A flexible and adaptable ontology for describing different data context (common knowledge or domain knowledge), event concepts (people, place, time) and objects collected by meaningful groups of different vocabularies is provided .
  • Data Visualization is enhanced and integrated through spatial and temporal mapping, filtering and linking system design.   

3. Data Semantically Enriched with Vocabularies and Knowledge Bases via Adaptable Mechanisms.
  • 18 international vocabularies used  for modeling common knowledge, and 5 domain specific vocabularies for place, time, art and humanity, or biology are applied.  3 knowledge bases like GeoNames, Wikidata, and Encyclopedia of Life are mapped and linked. 
  • The use of SPARQL language and endpoints  provide data analytic semantic queries both in local and external. In addition, data from the RDF triplestore can be easily used in 3rd-party applications.
  • Multiple DataClean Versions Mechanism:  we treat data cleaning as a kind of interpretation. Refined Versions (R Versions i.e. r1, r2, r3… ) provide different contexts to different needs of users. 
  • Multiple LinkedKnowledge Bases Mechanism:  more knowledge bases like DBpedia, WordCat, or LinkedGeoData can be linked in future via different R Versions without sacrificing  the  integrality of the original Version, encoded with DC 15 .
  • Multiple SemanticStructure Versions Mechanism:  different  interpretations results from the use of different vocabularies .  Co-exists of multiple R Versions  with different vocabularies or transforming vocabularies via SPARQL are  solutions. 
Reference:
  • Douglas, A. Vibert. "Forty minutes with Einstein." Journal of the Royal Astronomical Society of Canada 50 (1956): 99. P.100
  • Assante, M., Candela, L., Castelli, D., & Tani, A. (2016). Are scientific data repositories coping with research data publishing?. Data Science Journal, 15.. DOI: http://doi.org/10.5334/dsj-2016-006
  • Willis, C., Greenberg, J., & White, H. (2012). Analysis and synthesis of metadata goals for scientific data. Journal of the American Society for Information Science and Technology, 63(8), 1505-1520.
  • Cheng-Jen Lee, Andrea Wei-Ching Huang, Tyng-Ruey Chuang (2017) Metadata as linked data for research data repositories, International Symposium on Grids & Clouds (ISGC) 2017
  • 黃韋菁, 李承錱, 莊庭瑞 (2017) 結構資料的再次使用:語意、連結與實作, 圖書館學與資訊科學, 第 43 卷 第 1 期 2017 年 4月

Citation Information: Andrea Wei-Ching Huang (2017) Introduction: Metadata as Linked Data for Research Data Repositories. URL: https://andrea-index.blogspot.com/2017/03/metadata-as-linked-data-for-research.html

2014-09-15

Relations for Reusing (R4R) in a Shared Context: An Exploration on Research Publications and Cultural Objects

[[中文]]


Will the rich domain knowledge from research publications and the implicit cross-domain metadata of cultural objects be compliant with each other? A contextual framework is proposed as dynamic and relational in supporting three different contexts: Reusing, Publication and Curation, which are individually constructed but overlapped with major conceptual elements. A Relations for Reusing (R4R) ontology has been devised for modeling these overlapping conceptual components (Article, Data, Code, Provence, and License) for interlinking research outputs and cultural heritage data. In particular, packaging and citation relations are key for building up interpretations for dynamic contexts. Examples are provided for illustrating how the linking mechanism can be constructed and represented as a result to reveal the data linked in different contexts.



Conclusion


Responding to recent developments (Section 1) that have challenged research data, ar-chival and cultural heritage communities to come up with a contextual framework to support a dynamic and shared context environment, we have proposed a framework (Section 2) composed of three activity contexts that can be identified for a shared common understanding. In Section 3, the establishment of an ontology, Relations for Reusing (R4R) facilitates the representation of contextual links between resources in diverse contexts. Thus, a shared context between research and cultural heritage domains not only can be identified through three activity contexts for a common understanding, but rela-tions existing in different contexts can be established and represented through the R4R ontology. In Section 4, we used R4R to represent a use case from the Digital Archives Taiwan in different scenarios to show how linking data from these two domains can enhance the semantic relationships with each other, as well as increase the potential for reusing and remixing when both are contextually linked. The above discussions are the answers to the questions raised in Section 1, and we further discussed and presented a comparison of five existing relation ontologies that distinguishes the R4R from previous works in Section 5.


The advantage of designing a new conceptual model to describe relations in a shared context is to ensure that articles, datasets, software codes, provenance and license information can be treated as first-class contextual objects. At the same time, the module-like design of RRObject and RRPolicy can be practiced in isolation, and the unifying repre-sentation of their relations is semantically clear enough but not so structurally heavy-weighted that curators or researchers would find it difficult to apply. The contextual framework and R4R ontology can be applied to representing the interlinking relations of digital collections to a semantic web format, and help standardize this process. The au-thors of this work also plan in the future to explore more use cases to test the validity and effectiveness of the framework and the R4R ontology.

In sum, the daT(S010384) is a digital object with rich metadata descriptions that are curated in the Curation context. It is published as a cultural object Y, with unique iden-tification, and is cited as a science object Z, interpreted by the citation relation for addi-tional professional interpretations. At the same time, the citing research can benefit from the implicit information embedded in the institution’s cataloging vocabularies for more domain knowledge. Through the exploration of the Shared Context and R4R represen-tation, the daT(S010384) now is capable of moving from its traditional role and acting “as a citation of active knowledge”, as outlined in [22]. Creating knowledge out of interlinked data [23] is thus one step forward by packaging provenance and license for a policy-aware Reusing context. As a result, when data sharing does not need to remove the data's initial context but rather embed it in a shared context, the difficulty to interpret the reused data [24] may be expected to be reduced through the use of the contextual framework and the R4R ontology proposed in this study.

-----------------------------------------------------------


--------------------------------------------------------------

Reference (DOIs are auto generated by pdfx )

  • 1. Zimmermann, Andreas, Andreas Lorenz, and Reinhard Oppermann. An operational definition of context. Modeling and Using Context (2007): 558-571.  [DOI]
  • 2. Krafft, Dean B., et al. VIVO: Enabling national networking of scientists. Proceedings of the Web Science Conference. Vol. 2010.  [possible DOI]  [alternative DOI]
  • 3. Keßler, Carsten, Mathieu d'Aquin, and Stefan Dietze. Linked data for science and education. Semantic Web 4.1 (2013): 1-2.  [possible DOI]
  • 4. Haslhofer, Bernhard, and Antoine Isaac. data. europeana. eu: The europeana linked open data pilot. International Conference on Dublin Core and Metadata Applications. 2011.  [DOI]
  • 5. Malmsten, Martin. Making a library catalogue part of the semantic web. Proceedings of the 2008 International Conference on Dublin Core and Metadata Applications (2008): 146-152.  [possible DOI]  [alternative DOI]
  • 6. Ford, Kevin. LC Classification as linked data. Italian Journal of Library and Information Science, 4.1 (2013): 161.  [possible DOI]  [alternative DOI]
  • 7. Shotton, David. Semantic publishing: the coming revolution in scientific journal publishing. Learned Publishing 22.2 (2009): 85-94.  [DOI]
  • 8. Keivanloo, Iman, et al. Towards sharing source code facts using linked data. Proceedings of the 3rd International Workshop on Search-Driven Development: Users, Infrastructure, Tools, and Evaluation. ACM, 2011.  [DOI]
  • 9. Wendl, Michael C. H-index: however ranked, citations need context. Nature 449.7161 (2007): 403-403.  [DOI]
  • 10. Bechhofer, Sean, et al. Why linked data is not enough for scientists. Future Genera- tion Computer Systems 29.2 (2013): 599-611.  [DOI]
  • 11. Skinner, Julia. Metadata in Archival and Cultural Heritage Settings: A Review of the Literature. Journal of Library Metadata 14.1 (2014): 52-68.  [DOI]
  • 12. Courtright, Christina. Context in information behavior research. Annual Review of Information Science and Technology 41.1 (2007): 273-306.  [DOI]
  • 13. Peirce, Charles Sanders. “Elements of Logic”, Chapter 2: Division of Signs. In: C. Hartshorne and P. Weiss (eds.), Collected Papers of Charles Sanders Peirce (2) (Thoemmes Press, Bristol, 1998): 134–272  [DOI]
  • 14. Huang, Andrea Wei-Ching, and Tyng-Ruey Chuang. Social tagging, online commu- nication, and Peircean semiotics: a conceptual framework. Journal of Information Science 35.3 (2009): 340-357.  [possible DOI]
  • 15. Legg, Catherine. Peirce, meaning, and the Semantic Web. Semiotica 2013.193 (2013): 119-143.  [DOI]
  • 16. Beaudoin, Joan E. Context and its role in the digital preservation of cultural objects. D-Lib Magazine 18.11 (2012): 1.  [DOI]
  • 17. Seneviratne, Oshani, LalanaKagal, and Tim Berners-Lee. Policy-Aware Content Re- use on the Web. The Semantic Web - ISWC 2009 (2009): 553-568.  [DOI]
  • 18. Carata, Lucian, et al. A primer on provenance. Communications of the ACM 57.5 (2014): 52-60.  [DOI]
  • 19. Lagoze, Carl, et al. Fedora: an architecture for complex objects and their relationships. International Journal on Digital Libraries 6.2 (2006): 124-138.  [DOI]
  • 20. Yu, Chih-Hao, and Jane Hunter. Documenting and sharing comparative analyses of 3D digital museum artifacts through semantic web annotations. Journal on Compu- ting and Cultural Heritage (JOCCH) 6.4 (2013): 18:1-20.  [DOI]
  • 21. Gerber, Anna, and Jane Hunter. Authoring, editing and visualizing compound objects for literary scholarship. Journal of Digital Information 11.1 (2010).  [possible DOI]  [alternative DOI]
  • 22. Srinivasan, Ramesh, et al. Digital museums and diverse cultural knowledges: Mov- ing past the traditional catalog. The Information Society 25.4 (2009): 265-278.  [DOI]
  • 23. Auer, Sören, and Jens Lehmann. Creating knowledge out of interlinked data. Semantic Web 1.1 (2010): 97-104.  [possible DOI]  [alternative DOI]
  • 24. Borgman, Christine L. The conundrum of sharing research data. Journal of the American Society for Information Science and Technology 63.6 (2012): 1059-1078. [DOI]
  • 25. Associated data publication can be accessed at http://guava.iis.sinica.edu.tw/r4r/examples  [possible DOI]  [alternative DOI]

2008-12-17

ArticleRead (13): User Experience at Google: focus on the user and all else will follow

User Experience at Google: focus on the user and all else will follow. by Au, I., et al. (2008) In CHI 2008 Proceedings Extended Abstracts, ACM Press (2008), pp 3681-3686

Which research approaches should ensure that user experiences are interpreted to reflect underlined norms of online users, and promise a better identification for designers to predict user behaviors in the system design process? The case of Google in this article demon
strates a multi-method of user experience based on its corporate philosophy: “Follow the user and all else will follow”.

On the one hand, Google have traditionally sought to adopt their data-driven approach by applying web analytics of quantitative investigation in reflecting what is happening. On the other hand, built on the qualitative approach, Google interpret contextual factors of why users interact with the system designs via field research, diary studies, face-to-face interview. Such an approach is applied by the Google user experience (UX) team in exploring user behavior of Google Maps for the mobile application. They follow a method called Mediated Data Collection approach, in which participants and mobile technologies are assumed to mediate data collection about use in natural settings. Therefore, methods such as prior research on log analysis, recorded usage, focus group study, or field trial, telephone interviews, lab debriefs are combined to utilize the investigations on user behaviour.

This article stresses the bottom-up company culture as a key for designers and project managers to understand the essence of user experience. Three techniques are employed by: (1) injecting the corporate DNA to educate and train engineers and PMs about user experience (i.e. the ‘Life of a User’ training program and ‘Field Fridays’) (2) scaling to support hundreds of projects by UX team (3) helping focus projects on user needs by UX team or user research knowledge base.

Unlike the traditional desktop software design updated on annual basis, Google UX team practices some agile techniques to respond the rapid web cycles. For examples, solutions include guerilla usability testing, prototyping on the fly, online experimentation or enabling a live instant messaging dialogue between observers and moderator during lab-based testing.

The above three approaches are also combined with a global product perspective of designing for multiple countries. In sum, these 4 combinations of the Google case provide us an alternate analytic framework, and best enlist the methodologies for the studying of online user experience practically and implicitly.



2008-02-22

ArticleRead (3): A Definition of Information

A Definition of Information , By A.D. Madden, Aslib Proceedings vol. 52, No.9, p.343-, 2000.10.

In his article A.D. Madden has drawn some attention to the interpretation of information in the aspect of context.

After reviewing literatures defining information: as a representation of knowledge; as data in the environment; as part of the communication process; as a resource or commodity, the author has an attempt to further defining “information” in a perspective of “informing contexts”.

Three major elements in his Information-in-Context Model are defined as: “authorial context” which is a message being originated, “readership context” which is a message being received and interpreted, as well as “the message” which is the information being transmitted.

The idea of taking information reception and interpretation within personal and community paradigms in social-cultural contexts is valuable for most understating of the definition of information. However, the author rephrases the definition of information for the context-reliant model of information reception in the conclusion without clear explanation about “stimulus”, “system” and “system relationship” . The rephrases of the definition makes the information more blur in the end.

The general idea of Madden’s definition of information can be summarized as the figure shown.

Note: This review was mainly completed as a homework while taking the Humanity Informatics Class lectured by Professor
Ching-Chun Hsieh in December 2006.